Everyday IT · Troubleshooting & Escalation

Stop guessing. Narrow the problem.

Good troubleshooting is less about knowing every command and more about asking the right question in the right order.Scope it. Compare it. Test one layer. Make one change. Verify. Escalate with evidence when the blast radius gets bigger than the ticket.

The everyday troubleshooting framework

Use this before reaching for a random fix.

01

Scope the problem

Turn “nothing works” into a specific symptom and determine whether it affects one user, one device, one location, one resource, or everyone. Scope the problem

02

Use a known-good comparison

Change one variable at a time: same user on another device, another user on the same device, browser vs desktop, or another network. Use comparison deliberately

03

Test the simplest layer first

Check identity, connectivity, service state, permissions, and client state in a logical order. Do not rebuild a profile before proving the cloud service is healthy.

04

Ask what changed

Passwords, device swaps, updates, moves, permission changes, ISP issues, reboots, new software, maintenance, and whether the problem ever worked can all narrow the next test.

Make the smallest useful change

Discovery and remediation should be separate thoughts.

Before

Capture the current state

Know what was broken and what the important settings looked like before you change them. Screenshots, commands, timestamps, and notes make rollback and escalation easier.

Change

One reason, one action

Prefer a targeted change that tests your theory. If you change password, profile, DNS, permissions, and client settings at once, you may never know what fixed it.

Document

Leave the next tech evidence

Record the symptom, checks, relevant findings, action taken, and verification. Good notes are part of the fix.

When several things fail at once, look for the shared layer

Multiple devices or users developing the same symptom at the same time is evidence.

Multiple copiers

Look upstream before replacing hardware

If several scanners or copiers fail in the same way, inspect the common mail, DNS, authentication, network, or server dependency before treating them as separate device failures.

Missing drive

Find the real source first

A drive letter may point to SMB, while File Explorer may also show OneDrive or SharePoint content. Identify the resource and path

When to stop and escalate

Escalation is good judgment when the risk or scope changes.

Scope

Multiple users or business-wide impact

If the issue affects a department, site, server, tenant, or shared service, treat it differently from an isolated workstation ticket.

Privilege

The fix requires elevated or broad access

Domain Admin, Global Administrator, firewall, security-policy, or broad permission changes deserve review and authorization.

Data

There is a risk of loss or overwrite

Stop before deleting, resetting, mass-moving, restoring, unlinking, or overwriting data unless you understand the recovery path.

Security

The symptom may be compromise

Unexpected MFA prompts, suspicious sign-ins, mailbox rules, unknown forwarding, malware indicators, or privileged account anomalies should move into the security process.

Infrastructure

The change touches core systems

Domain controllers, DNS, DHCP, firewalls, mail connectors, Conditional Access, routing, backups, hypervisors, and production servers are not casual troubleshooting surfaces.

Uncertainty

You cannot explain the expected outcome

If you do not know what a change will affect or how to reverse it, stop and get another set of eyes.

Escalation is not failure. Uncontrolled guessing is.

When you do escalate, preserve the investigation. Escalate with evidence

Next Test

Choose the next step from the evidence you have, not from a generic checklist.

Risk unclear

Plan the change and rollback first

If the next action can affect identity, permissions, data, profiles, policy, or shared services, define the safe boundary before acting. Use change safety and rollback planning

Test succeeds

Verify the original workflow

Return to the reported symptom and prove the user can complete the required task. Verify before close

Test stops safely

Escalate with the investigation intact

Carry forward scope, comparisons, findings, failed tests, risk, and the exact unresolved question. Build the escalation packet

Security and infrastructure triage

Use these paths when the symptom crosses the ordinary endpoint boundary.

Security

Unexpected authentication activity

If the user did not initiate the MFA prompt or sign-in, preserve the evidence and approved controls. Triage a suspicious sign-in

Security

Endpoint threat alert

Confirm the detection, preserve evidence, use approved containment, verify the result, and escalate unresolved risk. Triage a malware alert

Infrastructure

Several users share the symptom

Stop repairing endpoints one at a time and identify the common dependency. Triage the shared service

Infrastructure

A production restart is proposed

Know the role, dependencies, recovery access, evidence, reason, and success criteria before restarting. Review restart safety