Skip to content
Avanet

Check Sophos Firewall SSD Health via SMART

A SMART value can be useful during hardware diagnostics. On physical XGS Appliances whose internal SSD is mounted as /dev/sda, a read-only smartctl query returns the endurance entry. For SFOS 22.0, however, Sophos does not document a general administrator command, a fixed drive path, or a cross-model wear threshold. Therefore, first verify the appliance, node, and device path, and treat the value as a diagnostic indicator rather than the sole basis for a replacement decision.

The safe path starts in WebAdmin: check usage and system behavior, generate a Consolidated troubleshooting report (CTR), and involve Sophos Support if you suspect a hardware issue. The query documented below reads existing SMART data and does not start a self-test. Other device paths, SMART tests, or repair commands should only be used in a specific support case.

⚠️ Important: The Advanced Shell provides direct system access. Do not try device paths, start SMART self-tests, alter partitions, delete files manually, or replace SSDs yourself. Even system fsck-on-nextboot in the Device Console must only be used when recommended by Sophos Support: the command helps with mount errors involving /sig, /conf, or /var, forces a file-system check of all partitions at the next restart, and can damage the file system if the hardware or SSD is unhealthy. on, off, and show enable, disable, or display the state; the default is off. In failsafe mode, SFOS may enable the check automatically, for example when the configuration, reporting, or signature database cannot start, a migration cannot be applied, or the deployment mode is missing. For the safe procedure, see Diagnose Sophos Firewall Failsafe Mode.

Safe diagnostic path

  1. In WebAdmin, go to Diagnostics > System graphs, select Disk usage, and choose a period that includes the start of the incident.
  2. Check whether reports, logs, quarantine, WebAdmin, or services are also affected. The disk usage display shows occupied storage space, not SSD wear or SMART health.
  3. Before a firmware upgrade, also review the firewall’s notices and notifications. SFOS 22.0 can report a required SSD firmware update for certain XGS Appliance models; in HA, each node is checked separately against the upgrade requirements.
  4. Under Diagnostics > Tools, for Consolidated troubleshooting report, select System snapshot and All log files, enter the reason for the diagnostics, select Generate, and then select Download. Debug mode is not required for the system snapshot.
  5. For I/O, file-system, boot, or recurring database errors, open a Sophos Support case and provide the CTR, timestamp, symptoms, and device details. Support decides whether additional shell or SMART diagnostics and a possible RMA are required.

For a capacity problem, see Check Sophos Firewall Storage and Manage Reports. Central Firewall Reporting can reduce reliance on local report data. A Sophos Firewall Health Check, by contrast, assesses configuration risks and does not replace hardware diagnostics.

What the visible signals mean

Disk usage is capacity, not wear

Under Diagnostics > System graphs > Disk usage, the x-axis shows minutes, hours, days, or months depending on the selected period, while the y-axis shows usage as a percentage. The legend separates signatures (orange), configuration files (purple), reports (green), and temporary storage (blue). A high value can affect reports and services, but it does not prove an SSD fault. Conversely, free capacity says nothing about hardware errors or remaining write endurance. The graph contains neither SMART attributes nor wear thresholds.

An SSD firmware notice is not a SMART finding

For some XGS Appliance models, an SSD firmware update intended to improve reliability may be mandatory before upgrading to SFOS 22.0 or later. A notification appears when action is required. This is a model-specific upgrade prerequisite, not a measured endurance value or automatic proof of a fault.

Before upgrading, therefore, verify that you have a current backup, sufficient storage space, a supported upgrade path, the release notes, and a maintenance window. For HA, both nodes must be reachable, healthy, synchronized, and independently meet the upgrade requirements; if either node does not meet them, it may block the upgrade. Start the upgrade only from the Primary Device.

SMART output is model-specific diagnostic data

SMART attributes, device names, and their meaning can differ by SSD, controller, appliance, and firmware.

Read endurance on an XGS Appliance with /dev/sda

Sign in to the firewall over SSH, open the Advanced Shell, and confirm that you are working on the correct appliance or HA node. Avanet ran this query on an XGS 3100 with SFOS 21.5.1 MR-1 Build 261. If the internal SSD is mounted there as /dev/sda, this command reads the full SMART data and displays only lines containing Endurance:

smartctl -x /dev/sda | grep Endurance
Sophos Firewall Advanced Shell showing smartctl output for the SSD endurance value
XGS 3100 with SFOS 21.5.1 MR-1: read-only smartctl query on /dev/sda

In Avanet’s example, the SSD reports a raw value of 1 for Percentage Used Endurance Indicator. On this drive, a low value represents little consumed write endurance. Do not apply this scale to other SSDs without verification, however: the attribute name, normalization, and raw value can be defined differently depending on the manufacturer and model. A value of 80 is therefore not a universal Sophos RMA threshold. Record how the value changes over time and consider symptoms, I/O errors, and the model-specific assessment.

The command and its model-specific interpretation have already been discussed publicly in practice because they are not documented in Sophos Firewall Help: Check Sophos XGS SSD lifetime on Administrator.de. The article recommends prompt replacement for a value above 80; Avanet deliberately does not adopt this community value as a universal Sophos RMA threshold. The screenshot above comes from an Avanet appliance, not from that article.

An SSD can fail without a prior SFOS warning. HA protects traffic, but not necessarily all locally stored data: logs, the mail queue, and quarantine can be affected on the failed node; the mail queue and quarantine are not synchronized between HA nodes. Central reporting, current backups, and a documented recovery procedure reduce the risk, but do not replace SSD monitoring.

If no line appears, either there is no matching Endurance attribute, the confirmed device path is incorrect for this appliance, or smartctl cannot read the SSD accordingly through the installed controller. Empty output proves neither health nor failure. Do not try other device names or run a SMART self-test speculatively.

For further interpretation:

  • Use /dev/sda only if this path has been confirmed for the specific appliance; do not guess NVMe or RAID paths.
  • Do not infer a general threshold from attribute names such as Endurance, Percentage Used, or Wear.
  • Do not treat a single value or a difference between two HA nodes as a replacement decision.
  • Do not interpret missing SMART output as either healthy or faulty.

For HA, run the confirmed command separately on both nodes because each node has its own SSD. Store the output together with the date, time zone, SFOS version, and node role. If there are symptoms, high or rapidly increasing values, or unclear attributes, also add the case number and the assessment from Sophos Support.

Document and validate the finding

A short, consistent history is more useful than one value:

  • Date and time with time zone: for example, 2026-09-05 10:30 CEST
  • Appliance and location: for example, XGS 2100 – HQ
  • Serial number and, for HA, role: Primary or Auxiliary
  • SFOS version and build
  • Symptom: storage warning, I/O error, boot error, report or database problem
  • Disk usage period and visible change
  • CTR file and support case number
  • Practical SMART query: confirmed device path, exact command, and unmodified output

A normal graph or a single SMART value does not complete the diagnosis. As validation, check whether storage usage and affected functions remain stable after the safe corrective action and whether errors return during the agreed observation period. For virtual firewalls, drive condition, datastore, I/O latency, and errors primarily belong in the monitoring of the hypervisor and storage platform.

For suspected temperature or fan issues, see Check Temperature and Fan via SSH; SNMP Hardware Monitoring can add status and trends for supported hardware sensors. Neither check replaces SSD diagnostics by Sophos.

If the appliance is unstable or unreachable

Do not run restart, file-system, or repair commands speculatively. Record the last known working time, changes before the incident, LED state, and accessibility over HTTPS, SSH, and the serial console. If there is a complete loss of power, first test another power outlet and power cable; for dual-PSU appliances, test the second input and, on models designed for it, another hot-swappable power supply. A photo or video of the LED and startup behavior can speed up the RMA assessment.

For an appliance that does not boot, test HTTPS and SSH over LAN and WAN, as well as the serial console directly through DB-9, a serial-to-USB converter, or the Micro-USB console port available on newer XGS Appliance models. In Device Manager on the administrator’s computer, check for driver or connection errors. Set 38400 baud, check the state at several intervals, and capture visible errors in screenshots. If the console remains silent, cross-check with a second cable or computer. Sophos decides whether a device qualifies as DOA only after its own assessment. For the complete internal procedure, see Sophos Hardware Failure: Prepare RMA and Replacement.

If the appliance remains reachable, reproduce the error immediately before collecting the data and note the exact time with the time zone. Under Diagnostics > Tools > Consolidated troubleshooting report, select System snapshot and All log files, enter the reason, select Generate, and then select Download. Upload the encrypted report to the support case. Some CTR logs contain only the number of lines configured through the CLI; for an older event, therefore, also save the affected Troubleshooting logs individually. Debug mode is off by default and is not required for the system snapshot; debug logs increase storage usage and must be disabled again after a targeted recording. In HA, logs and reports are not synchronized and must be collected and clearly assigned per node. For the complete procedure, see Save Sophos Firewall Logs for Support and Analysis.

If Sophos requests remote access for diagnostics, you can generate a time-limited Access ID under Diagnostics > Support access. The firewall establishes a secure outbound control connection over TCP 22 to *.apu.sophos.com; an upstream router must allow it. Enable Support access, confirm with OK, select the duration, select Apply, confirm again with OK, and copy the unique ID under Access status. Share the Access ID only in the support case. It gives Sophos access to WebAdmin and the shell without the administrator password; inactive sessions end after 15 minutes. You can disable access at any time, and it is disabled after the case. For the detailed internal procedure, see Enable Sophos Firewall Support Access for Avanet.

Prepare Support and RMA

Have the following ready before escalating:

  • exact issue description, start time, frequency, and impact;
  • model, revision, serial number, SFOS version, and build;
  • HA status and affected node;
  • disk usage history, relevant error messages, and CTR;
  • a current downloaded configuration backup;
  • keep the backup encryption password and matching Secure Storage Master Key securely available, but share them only through the secure method specified by Sophos;
  • license and support status.

A SMART value does not automatically trigger an RMA. According to Sophos, the RMA process starts with fault identification and device information, followed by a support case and validation; Sophos may request further diagnostics. Model, revision, firmware version, serial number, and HA membership belong in the RMA form.

For HA, also determine which node will be replaced. The rebuild described here applies only to active-passive, not active-active, and causes downtime. Beforehand, record the model, revision, initial Primary, and firmware version and build of both appliances with system diagnostics show version-info.

  • Prepare the replacement appliance: If the identical firmware build is unavailable, request it from Sophos Support. Connect a DHCP client to Port 1 and open https://172.16.16.16:4444. Configure Port 2 through the Setup Assistant only for WAN and internet access; do not initially configure any other interfaces. After reimaging or updating, verify the build again with system diagnostics show version-info.
  • Replace the Auxiliary: The healthy Primary temporarily operates alone. Bring the replacement appliance to the same firmware version and build, claim it in Central, and transfer the license from the failed Auxiliary. Disable HA on the healthy Primary, use service -S | grep msync to confirm the UNTOUCHED or STOPPED state, reconnect the cables, and rebuild active-passive HA with the healthy appliance as Primary.
  • Replace the Primary: Deregister the healthy Auxiliary from Central and confirm under My Products > Firewall Management > Firewalls that it is no longer listed. Save its current backup. Bring the replacement appliance to the same firmware version and build, claim it, transfer the license, restore the backup, and let it take over traffic as a standalone appliance after reconnecting the cables. Then reset the healthy former Auxiliary to factory settings, claim it again, and reconfigure active-passive HA with the replacement appliance as Primary.

The replacement provides working hardware, but does not guarantee operational readiness. Document the license transfer, backup compatibility, Central assignment, cabling, and functional test before the maintenance window. For an overview of the roles, see Sophos Firewall HA Cluster: Active-Passive, Active-Active, and Auxiliary Appliance.

How Long Do I Get Warranty on Sophos Hardware? summarizes warranty and support basics. If Sophos requires a reimage, follow Reinstall Sophos Firewall OS: Reimage with USB Stick.

Final checks

  • Checked Disk usage and its time period without treating capacity as SSD health.
  • Recorded symptoms, timeline, model, serial number, SFOS version, and HA node.
  • Saved a current backup outside the appliance and kept the recovery secrets available.
  • Generated a CTR with System snapshot and All log files and stored it securely.
  • Did not use guessed device paths or run SMART self-tests, fsck, deletion, or repair commands without approval.
  • Opened a support case for suspected hardware; additional diagnostics only follow specific Support instructions.
  • Planned an RMA or hardware replacement only after Sophos validation.

FAQ

Can I see SSD health directly in WebAdmin?

No. Disk usage shows occupied space, not SSD wear or a SMART health status. Firmware notifications may report a required SSD firmware update, but they are not a wear value either.

Which SMART command and device path should I use?

If the internal SSD of the physical XGS Appliance has been confirmed at /dev/sda, smartctl -x /dev/sda | grep Endurance reads the endurance entry without starting a self-test. Sophos does not publish a universal administrator command for other models or device paths; do not guess paths.

At what SMART value must the SSD be replaced?

The public SFOS 22.0 documentation provides no universal RMA threshold. Support evaluates the model, diagnostic output, symptoms, and other device data; one value determines neither fault, warranty, nor RMA.

Is an unremarkable SMART value an all-clear?

No. Boot, I/O, file-system, database, or report issues still require investigation. Likewise, high disk usage does not prove an SSD fault.

How do I check the SSD in an HA cluster?

Record upgrade requirements and symptoms for each node. If /dev/sda has been confirmed, run the read-only endurance query separately on both nodes and record the node role, time, and output. Do not guess other paths or tests; interpret differences only with a model-specific assessment.

Does this apply to virtual Sophos Firewalls?

The SFOS storage graph remains a capacity signal. Assess the virtual disk and underlying medium with hypervisor and storage monitoring, not as though they were the SSD in a hardware appliance.