Sophos Firewall in failsafe mode: proceed safely
When a Sophos Firewall reports failsafe, treat it as a recovery incident. The public SFOS 22 documentation doesn’t describe a general diagnostic command that reliably identifies the cause in every failsafe condition. This runbook therefore deliberately uses neither show failure-reason nor an invented replacement.
First preserve the displayed message unchanged, then assess the platform, HA role, firmware build, and most recent change. Don’t manually modify databases, signatures, rules, or system files while the cause remains unproven.
⚠️ Document before restarting: A restart can change the visible error condition. Photograph the complete console, including the first boot messages, and record the time, model, serial number, full SFOS version with build, and affected Node for HA. Reset to Factory Defaults, Remove Firewall Rules, and manual database or file-system changes aren’t safe initial diagnostics.
Capture the failsafe message safely
- Use the available direct console: local or serial for hardware, or the hypervisor console for a VM. This path remains available when WebAdmin or the network can’t be reached.
- Photograph the message verbatim with the boot lines immediately above it. Don’t substitute sample output from the internet for the actual message.
- Record the platform and model, serial number, full SFOS build, time of failure, and last known working time.
- Record recent changes, especially firmware upgrades, rollbacks, restores, rule or object changes, VM resources, virtual disks, vNICs, and storage events.
- For HA, record the affected Node, role, peer status, and most recent role change. Don’t restart both Nodes together or disable HA on suspicion.
WebAdmin being unreachable on its own doesn’t prove a failsafe condition. If the firewall starts normally and traffic still flows but only the interface is unresponsive, first perform the targeted check or restart the WebAdmin GUI. An explicitly displayed failsafe message, however, is handled as a startup or recovery problem.
Narrow down documented causes
Sophos documents several specific but non-exhaustive triggers for SFOS 22. The release notes for SFOS 22.0 MR2 Build 546 include a failsafe condition caused by a full configuration partition (NC-181331), a logging daemon that didn’t start on the primary HA device (NC-180110), failsafe on the initial primary device after upgrading to 22.0 GA (NC-177441), and Failed to start Red server service (NC-178906). These entries establish known fixed defects, not a general repair matrix.
If the observed condition, platform or HA role, and version history precisely match a release-note entry, include the issue ID and installed build in the support case. A partial match proves neither the same cause nor that an uncoordinated firmware change will recover the firewall. Sophos quotes the displayed message only for NC-178906. If the firewall starts normally and only a RED tunnel is offline, use RED troubleshooting instead.
Software appliance
For an SFOS 22 software appliance installed on your own hardware, Sophos specifies, among others, these minimum requirements relevant to the running system and explicitly states that the firewall enters fail-safe mode if they aren’t met:
- x86-64 CPU and Legacy BIOS
- at least
4 GBRAM - at least
32 GBHDD or SSD;64 GBis recommended 2network interface cards
Document the current state before making changes. For a confirmed mismatch, ask Sophos Support whether a supported resource adjustment is sufficient or a controlled reimage is required. Don’t improvise disk or boot-mode changes. Platform differences and resource requirements explains hardware, virtual, and software appliances in more detail.
Virtual and hardware appliance
For a VM, record the existing and connected vNICs, virtual disks, CPU, RAM, disk controller, and recent hypervisor changes. Cloud and hypervisor platforms have their own supported requirements; don’t automatically apply the software-appliance values to every virtual deployment.
For hardware, include recurring boot failures, power events, I/O or SSD indications, temperature, and fans in the support case. One successful restart doesn’t rule out a defect. The existing checks for temperature and fans, SSD health, and RMA preparation help prepare the case.
Preserve logs and backup
If Advanced Shell remains accessible, sysinit.log, syslog.log, and, for a database indication, postgres.log are relevant official troubleshooting logs. For the documented RED message, also include red.log. This runbook deliberately provides no shell commands for copying, deleting, or repairing: the access path and available tools can differ in a recovery condition.
When WebAdmin becomes available again, generate a troubleshooting archive or CTR. The procedure is described in Preserve Sophos Firewall logs for support. Logs and archives can contain confidential network, user, and configuration data and must only be transferred securely to authorized recipients.
Before a restore or reimage, a suitable backup, its password, and the associated Secure Storage Master Key must be available. Sophos Firewall backup and restore explains these dependencies.
Use fsck-on-nextboot only with Sophos Support
system fsck-on-nextboot isn’t a general health check. Sophos warns that it must only be used when recommended by Sophos Support. It’s intended for mount errors on /sig, /conf, or /var; if the hardware or SSD isn’t healthy, the check can damage the file system.
The documented Device Console syntax is:
system fsck-on-nextboot [on | off | show]
on forces all partitions to be checked at the next device restart, off cancels it before that restart, and show displays the current configuration; the default is off. SFOS can schedule the check automatically in failsafe mode when the configuration, report, or signature database can’t start, a migration can’t be applied, or the deployment mode can’t be found.
Don’t restart solely because on is displayed. First include the support recommendation, stable console access, power supply, backup, possible I/O errors, and, for HA, peer state in the same maintenance plan. Sophos specifies neither a fixed duration nor a guarantee of success.
Choose a safe recovery path
- Software requirement clearly below minimum: Preserve the current state and perform the adjustment or reimage path confirmed by Sophos Support in a maintenance window. Observe the next startup through the console.
- Message matches a known release-note issue: Preserve the full build and issue ID. Have Sophos Support confirm the recovery, upgrade, or rollback path for this condition.
- HA incident: Protect the peer that’s still operating. Don’t perform simultaneous restarts or unplanned role or cluster changes. See Sophos Firewall High Availability for the fundamentals.
- Storage, database, rule, NPU, I/O, or unknown error: Don’t remove files or configuration elements on suspicion. Escalate to Sophos Support with the console evidence, logs, and timeline.
- Reimage: Use only when Sophos Support or a documented recovery plan justifies it and all restore prerequisites are available. See Reinstall Sophos Firewall OS for the controlled procedure.
After any action, don’t check only WebAdmin and a successful login. Verify normal console startup, HA status, interfaces, routing, and the required internet and VPN connections. If the message returns, add the time and unchanged output instead of making more spontaneous changes.
Escalate to Sophos Support
Escalate a failsafe incident as soon as the cause can’t be confined to a documented resource mismatch that can be corrected safely. The Sophos Support case should include at least:
- a photo or copy of the complete failsafe and boot messages
- model, serial number, platform, and full SFOS build
- time of failure, last working time, and change timeline
- for HA, Node, role, peer status, and most recent role change
- relevant logs or CTR and any storage, I/O, NPU, or power indications
- available backup and the status of its password and SSMK, without disclosing secrets in the ticket text
- restarts or changes already performed and their results
If WebAdmin is available and Sophos requests remote access, time-limited Support access can be turned on under Diagnostics > Support access. Share the generated access ID only through the agreed support channel; access can be turned off at any time.