Diagnose Sophos Switch: preserve logs and support data
Good switch diagnostics do not begin with a restart or maximum logging. This runbook moves from the observed symptom through the smallest practical test to usable evidence. It does not cover a full redesign of VLANs, routing, or network architecture.
Quick workflow
- Define the symptom and success criterion with the table below.
- Check the switch identity and baseline state.
- Preserve symptom-relevant counters before changing anything.
- Run the smallest suitable diagnostic test.
- Download RAM and flash logs; raise the log level temporarily only if required.
- Compare the result with the success criterion or a known-good port.
- Revert temporary logging and Remote Assistance, then verify the intended state.
- If unresolved, open a support case or attach the redacted evidence package to the existing case.
Define the symptom first
Record the following once, before the first intervention:
- Timing: first and last observation with date and time zone; constant or intermittent;
- Impact: complete outage, packet loss, low throughput, missing PoE power, missing Fusion synchronization, or only a warning;
- Scope: switch serial number, site, port, VLAN, connected device, and affected users or services;
- Preceding change: configuration, cabling, firmware, power, or topology change and its time;
- Reproduction: exact source, destination, expected and actual result, and frequency;
- Workaround: whether another port, path, or power source works;
- Business impact: affected sites and services and available redundancy.
Do not round times from memory later. Exact times connect user observations, port counters, event logs, Fusion communication, and support data.
Map the symptom to the first check
| Symptom | Check first | Suitable view or function |
|---|---|---|
| Switch or clients completely unreachable | Fusion status, last contact, power, and management path | System details, alerts; local UI only if reachable |
| One port loses packets or link | Its RX/TX counters and errors, peer, and cable | Port Statistics, then targeted Cable Diagnostics |
| Wrong device or VLAN suspected on a port | Learned MAC address, port, and VLAN | MAC Address Table |
| PoE device does not start or drops out | Total budget and per-port current, voltage, and power | PoE Power Usage |
| Switch cannot reach a target network | Local source, layer-3 path, and return path | Network Diagnostics: ping, then traceroute if needed |
| SFP path looks abnormal | Installed module details and capabilities | SFP Module Info |
| Fusion task fails or switch reconnects | Alert, last contact, task state, time configuration, and agent communication | System details, System time, Task queue, Sophos error reporting |
| CPU or memory load suspected | Current utilization during the symptom | Resource Usage |
One measurement does not prove a cause. A successful ping, for example, proves reachability only for that time and path. Assess counters over a defined interval or against a known-good port.
Confirm access and permissions
For central diagnostics, the Sophos Fusion account must reach the correct Tenant—the relevant customer environment—and My Products > Switches. Read access is enough to capture visible state; changing Log settings, Remote assistance, or other settings requires suitable write access. If a button is absent, check tenant, device assignment, and administrator role before broadening permissions.
Reports under Diagnostics open the local admin console. The admin client must reach the switch in the same subnet. Local sign-in uses a separate local switch account, not automatically the Sophos ID or Fusion role.
Check the baseline in Sophos Fusion
Open My Products > Switches > Switches and capture only the mandatory set:
- Serial no., Model, Name, MAC Address, and firmware version;
- State, last-event time, and recent communication events in System details;
- number and content of Alerts;
- open or failed entries in the Task queue, the queue of Fusion jobs;
- SNTP status, primary and secondary NTP servers and ports, Timezone, and daylight-saving settings under System time;
- Configuration source, showing whether the setting is managed locally or centrally.
Add only symptom-specific fields: Powered on for restarts, Connection usage for connectivity faults, site and tags for assignment, Parent Site or Parent Stack as the higher-level site or stack assignment, and available/delivered PoE budget for power faults.
Correlate timestamps only when the time zone, daylight-saving rule, and observed offset are known. System time does not prove a visible current switch clock. Obtain current switch time only from a timestamped switch log/event or a model- and firmware-specific local view verified independently. Document an offset and correct time settings only after the first log capture, if change control permits.
Interpret Fusion states
| State | Next step |
|---|---|
| Waiting for sync | Check Task queue and connectivity; do not stack more changes. |
| Pending | Check task details and preceding jobs. |
| Syncing | Wait; do not start a parallel local write. |
| Out of sync | Capture configuration source, local differences, and task errors. |
| Suspended | Move outdated firmware into the controlled firmware-maintenance process. |
| Manual synchronization needed | Preserve cause and Task queue; do not use Reapply all settings as the first diagnostic step. |
Fusion shows locally configured values only when synchronized or replicated. A clean central view therefore does not rule out local drift. Record page, field, value, and Configuration source together, then follow only the branch that matches the symptom.
Use diagnostic views selectively
Links under Diagnostics open local reports in a new window. If only that launch fails, check the same-subnet requirement before inferring a switch outage. Open local switch management uses the same route.
Resources, ports, and address table
- Resource Usage opens Monitor > Realtime Meters. Compare CPU and memory during the symptom with a normal period; one value remains an observation, not proof of a cause.
- MAC Address Table opens Monitor > Dynamic MAC Address and Monitor > Static MAC Address. If the assignment differs from the expected port or VLAN, verify the peer and intended path first; otherwise continue with the next symptom-relevant branch.
- Port Statistics opens Monitor > Statistics > Ports, with inbound/outbound packet counts and per-port TX/RX errors.
Capture the affected and a working comparison port locally. In central management, also preserve RX discard and TX discard under Statistics > Port. Use discard values from a local view only if independently verified for that model and firmware. Repeat after a short interval and compare changes over the same window; do not present cumulative counters as a current rate.
If errors or discards rise only on the affected port, narrow down cable, peer, MAC, and VLAN there. If they remain unchanged, do not infer a bad port; follow the next path matching the impact. Central Statistics also provides L2, L3, 802.1X security, Port, and RMON, including invalid BPDUs, discarded DHCP messages, or CRC alignment errors.
PoE, cable, and SFP
- PoE Power Usage opens Monitor > Dashboard > PoE Power Settings. Compare total budget and affected-port current, voltage, and power with an expected or working device.
- Cable Diagnostics opens Analyze > Diag Tools. Select only the known target port and choose Test.
- SFP Module Info opens Monitor > SFP Module Information and shows the details and capabilities reported by the module.
Before a cable test, determine whether it can affect production or your management path, and preserve counters and link state. Record port, peer, cable run, result, and test time. Isolate the cable, peer, or module only when that path differs; do not generalize one result to other ports. If observations remain unchanged, continue with the next symptom-relevant branch.
Ping and traceroute
Network Diagnostics opens Analyze > Ping Test and Analyze > Trace Route. Ping tests reachability from the switch; traceroute observes the layer-3 path. Record target IP or resolved name, expected path/reachability, start time, and complete result including loss or stopping point.
Test a known-reachable target on the relevant path first, then the failing target. If the comparison also fails, inspect the common local path. If only the failing target differs, continue at that divergence. Do not use arbitrary internet hosts as proof of an internal VLAN or routing fault.
Preserve events and logs
Event Logging opens Monitor > Local Logging for event selection and Monitor > Log Table for entries. Diagnostics > RAM logs and Diagnostics > Flash logs provide Download.
- RAM logs contain recent entries and are lost on power-off or restart. Download them immediately for a current or freshly reproduced fault.
- Flash logs survive power-off and restart, but verbose logging increases writes.
Full RAM logs and Flash logs overwrite their oldest entries, so use Download to preserve both stores before more reproduction attempts.
Change log levels in a controlled way
Under Diagnostics > Log settings, custom settings can be enabled or disabled; Not set uses the local setting. Severity from highest to lowest is Emergency (0), Alert (1), Critical (2), Error (3), Warning (4), Notice (5), Info (6), and Debug (7). A level includes all higher severities. Flash defaults to Critical.
Only if existing logs omit the required event:
- Record custom Log settings, RAM log level, and Flash log level, then download existing logs.
- Define a short reproduction window, stop condition, and owner.
- Select only the required level, save with Update, and note activation time. Keep Debug, especially for flash, as short as possible.
- Reproduce once under control and download both logs again.
- Immediately restore the recorded values, select Update, reload the page, and verify level and On/Off/Not-set state.
Stop and restore values if impact grows, management is lost, or service interruption is unplanned. Never leave verbose flash logging enabled; the event volume can cause excessive device wear.
Support functions with side effects
Sophos error reporting and Remote Assistance
Sophos error reporting is enabled by default. It sends agent logs to Fusion for firmware-upgrade or backup failures, disconnect/reconnect events, and failed task synchronization. Its description says these concern switch–Fusion communication and contain no configuration or network data.
Enable Remote assistance under Diagnostics only for an existing case. Record case number, purpose, approval, and contact; select the shortest suitable duration; click Activate; and note start and expiry in the case. After the session, choose Deactivate and verify the displayed state.
Only when Support instructs you
These actions alter diagnostic state and are not general repair steps:
- Take a switch snapshot runs system commands; output appears in the Task queue.
- Restart Sophos Fusion agent restarts Fusion agent processes on the switch.
- Clear core files deletes core files created when processes stop responding, freeing space.
Before each action record the case, instruction, time, prior state, and expected result. Before Clear core files, confirm that Support no longer needs this evidence.
Build and redact the evidence package
Keep originals unchanged in an access-restricted folder. Share a copy and log all edits. The complete handoff contains:
- Case overview: fault, current impact, start and time zone, reproduction, expected/actual result, workaround, and business impact.
- Device identity: model, serial, firmware, Fusion agent version, site, parent, management state, and uptime.
- Time basis: SNTP, NTP servers/ports, time zone, daylight saving, configuration source, and known offset; current switch time only from a timestamped log/event or independently verified local view.
- State evidence: alerts, last contact, task details, affected ports, VLANs, and configuration source.
- Tests and measurements: target, time, full result, before/after counters, and symptom-specific ping, traceroute, cable, PoE, SFP, or resource data.
- Logs: RAM and flash downloads before/after reproduction, with filename and capture time.
- Change log: every diagnostic change, start/end, result, and verified rollback; clearly flag active access such as Remote Assistance and its expiry.
- Assessment: separate observed facts from interpretation; identify redacted or pseudonymized content.
Remove passwords, session tokens, API bearer tokens, private keys, SNMP communities, and unrelated secrets from copies; revoke compromised values. Minimize personal and irrelevant business data.
IP/MAC addresses, ports, and device names may be needed for path correlation. Use stable placeholders such as CLIENT-A and SWITCH-UPLINK-1 throughout. Preserve timestamps, time zone, error codes, counters, and relationships. Retain the unchanged original internally.
Escalate to Sophos Support
Personal support and Advanced RMA require a suitable Switch Support and Services subscription per deployed switch. Check status before escalation; see How is Sophos Fusion licensed?. Local web UI and CLI remain separate management routes. A Sophos ID with customer/Tenant access is required.
Support covers installation, administration, operation, behavior contrary to documentation, and general configuration questions; implementation or a new deployment is outside normal scope. Open a technical case with Sophos Support Assistant:
- Sign in with the Sophos ID, describe the issue, and answer diagnostic questions.
- Ask the Assistant to create a case if needed.
- Review and submit. The case exists only when a case number appears.
- Keep the evidence package and later reproductions under that number.
Evaluate suggested changes against maintenance window, impact, and rollback. Do not open parallel cases for one event. Add new reproductions, increased impact, and evidence with exact times to the existing case. Support Portal, chat, and phone are alternatives; verify regional details in Sophos contacts.
Boundary between support case and RMA
An unreachable switch, faulty port, or repeated restart is not yet a confirmed hardware defect. First document power, cabling, peer, firmware/sync state, and logs. Sophos Support decides whether to open a replacement.
Do not reset, open, discard, or ship the switch unexpectedly. Factory Reset and deleting core files can destroy evidence. Once Support confirms a defect, transfer serial, model, support status, delivery address, privacy risk, and written return instructions into Prepare a Sophos hardware defect and RMA, where warranty, support coverage, and Advance Hardware Replacement are assessed separately.
Final check and safe rollback
Verify only the actual operational and diagnostic state:
- temporary RAM/flash levels and custom Log settings match their documented prior state;
- Remote assistance is off or approved expiry is confirmed, and Sophos error reporting has been checked against prior state;
- management, ports, PoE devices, and uplinks are in the expected state;
- the Task queue has no untreated diagnostic error and configuration source was not changed unintentionally;
- the success criterion was retested from user and network perspectives;
- if unresolved, the observation window and next escalation point are defined.
“Not reproducible” is not a successful fix. If a temporary change cannot be fully reverted, document the impact, notify the change owner, and continue through the independent management path or agreed recovery procedure.