Check Sophos Firewall daily: operations checklist for admins
A Sophos Firewall can remain technically reachable while showing early warning signs: a WAN connection flaps, disk usage grows, a service reports an error, or failed admin sign-ins increase. A short, repeatable operations check makes these changes visible before they become a longer outage or a security incident.
This workflow covers current operations. The Sophos Firewall Health Check, by contrast, determines whether selected settings comply with Sophos and CIS recommendations. The two checks complement each other but do not replace one another.
The ten-minute check
A fixed workflow is enough for the daily overview:
- In the Control Center, record the model, firmware version, and build, open new messages, and click the status icons for services, WAN, interfaces, and VPN.
- Under Diagnostics > System graphs, compare CPU, memory, load average, disk, and important interfaces with the normal baseline.
- Check the security dashboards and the reports in use for the latest fully available period for new IPS, web, application, zero-day, or Active Threat Response events.
- Open the Log Viewer in the upper-right corner, select the module and time range, and check failed administrative sign-ins, unusual sources, and the associated services.
- For HA, take account of the processing node and, for central reporting, the expected data source.
- Document every relevant deviation with time, firmware, node, source, affected service, and next step.
- Do not restart services, delete logs, or broaden rules because of a single spike. Correlate the trend, logs, and actual function first.
This short check is not intended to trigger a configuration change every morning. Its value lies in identifying changes early and deciding clearly whether observation, diagnosis, or escalation is required.
Four views, four different statements
The most important views do not show the same thing:
- Control Center: current overview of system, services, WAN, interfaces, VPN, uptime, and actionable messages.
- System graphs: timeline of CPU, memory, load average, disk, WAN transfer, and interface counters.
- Reports: consolidated analysis of a completed period. Some widgets and report data aren’t updated in real time.
- Log Viewer: individual events with time, module, action, source and destination information, and, depending on the log type, Rule ID or other details.
A red widget is a signal, not yet a complete diagnosis. Likewise, a current green status doesn’t prove that no brief error occurred overnight. Only the combination of current state, timeline, report, and individual event produces a reliable picture.
Check system state and availability
Read the Control Center first
In the Control Center, start with new messages. Sophos shows registration, licensing, reporting, WAN, or upgrade issues there, among other things. Some messages disappear automatically after the issue is resolved and can’t simply be deleted manually. Every relevant message therefore needs an owner and a traceable next step.
Always read the color together with the details. For Services, Warning means at least one service has stopped, while Alert means at least one service failed to start. For WAN and VPN, Warning means that up to half of the configured connections are down, and Alert means that more than half are down. These counts don’t account for business importance. One failed primary tunnel may therefore be more urgent than several intentionally inactive connections. Click the corresponding icon to see the affected entries.
Next, compare the services, WAN connections, interfaces, and VPNs actually in use with the expected state. A red interface isn’t automatically an outage: an unused port without an IP address or a physical parent interface for a VLAN can be expected to appear red. What matters is a deviation from the documented design.
Treat the Messages widget as an action queue, not a general event feed. In particular, create the Secure Storage Master Key when prompted so sensitive values such as passwords receive the additional protection. A WAN-access message means that WebAdmin (HTTPS) and CLI (SSH) are reachable from the WAN zone: if remote administration is required, use a VPN or a narrowly scoped Local Service ACL Exception for specific management hosts or networks rather than leaving broad WAN access. For a reports-disk message, reduce usage below the lower threshold; merely dropping below the higher threshold isn’t sufficient. Messages that depend on a requirement disappear after it is met and can’t be deleted manually.
Read Active threat response by feed and action. MDR and Sophos X-Ops show blocked-threat counts, NDR Essentials shows monitored threats, and third-party feeds show synchronization state as well as blocked threats. Configure leads to protection setup, Reports to the supporting report, and More details expands the widget. The Reports action isn’t available on models without local reporting. A nonzero monitored count isn’t the same statement as a blocked count, and a third-party synchronization problem must be separated from the threat count.
The Reports widget is a shortcut to up to five critical reports selected according to the subscribed modules, not a complete live event list. The mapping is:
- High-risk applications — Web Protection
- Objectionable websites — Web Protection
- Web users — Web Protection
- Intrusion attacks — Network Protection
- Web server protection — Web Server Protection
- Email usage — Email Protection
- Email protection — Email Protection
- Traffic dashboard — Web Protection or Network Protection
- Security dashboard — Web Protection or Network Protection
High-risk applications, Objectionable websites, Intrusion attacks, Web server protection, and Email protection refer to yesterday; Web users ranks the top ten users by web bytes transferred yesterday. Email usage shows transferred email bytes, while Traffic dashboard and Security dashboard summarize traffic categories and denied activity. A missing tile can therefore reflect the subscription rather than zero activity. Click the report name to inspect it or the download icon to preserve it. For the separate time ranges and drill-downs behind endpoint, user, zero-day, TLS, and session signals, continue with How to interpret User & Device Insights.
Traffic insight summarizes traffic processed during the last 24 hours. Web activity shows the transfer trend plus average and maximum bytes; cloud applications show detected apps and bytes in and out, with hover details for New, Sanctioned, Unsanctioned, and Tolerated states. The remaining graphs rank the top five allowed application and web categories by bytes, blocked application categories by hits, and hosts denied network access for health reasons. Click a cloud graph or category bar to open the corresponding cloud-application page or filtered report; use that drill-down before treating a peak or top-five entry as an incident.
A daily check includes at least:
- unexpectedly stopped or degraded services;
- WAN links that are down or repeatedly change state;
- production interfaces with new errors, drops, or collisions;
- important VPN connections that are disconnected contrary to the operations plan;
- an unexpected restart or unusually short uptime;
- new messages that don’t yet have an owner or ticket.
Read System graphs against a baseline
Under Diagnostics > System graphs, look for patterns rather than just individual spikes. Evaluate CPU, memory, and load average together with the core count, traffic, and affected period. A short spike during backup, reporting, or a pattern update has a different meaning from continuously high load during normal traffic.
For disk usage, the trend matters most. One period of high utilization and steady growth are different problem patterns. For interfaces, traffic, errors, drops, and collisions help distinguish firewall load from a link, duplex, cable, or switch problem.
For comparable evidence, record the graph type and time range and use the same time range in the ticket. Interface graphs only show a separate graph for VLANs in the WAN zone. SFOS combines VLANs in other zones with the graph for their physical parent interface. You therefore can’t expect a separate, uneventful LAN VLAN graph when SFOS doesn’t provide one.
The detailed explanation of load average, offloading, TLS inspection, and System graphs is available in Understanding Sophos Firewall performance specifications. For storage limits and on-box reporting, see Checking Sophos Firewall storage and reports.
⚠️ A single high measurement isn’t yet a reason to restart a service. Time, duration, recurring pattern, affected traffic, and logs must correlate first. Before a restart, preserve the relevant logs and, for an incident, a CTR.
Check security events and admin sign-ins
Read reports for changes
The daily security review focuses on new or significantly changed patterns. Depending on the enabled features, the following areas are particularly relevant:
- Reports > Dashboards > Security dashboard for the consolidated overview;
- Reports > Network & threats > Intrusion attacks for IPS events;
- Reports > Network & threats > Active threat response for blocked IoCs;
- Reports > Applications & web for risky, unwanted, or blocked web and application use;
- zero-day, Security Heartbeat, or wireless reports if these features are used in production.
In the selected report, first set the required date range and then click Generate. Use Filter to narrow the results to the relevant source, action, or rule. The available download formats preserve the displayed data as ticket evidence. Record the time range and time zone with the export so that a later comparison doesn’t use two different windows.
Not every event is an incident. Source, destination, user, rule, action, frequency, and timing are decisive. A single access attempt from a country doesn’t justify blocking the entire country. Repeated attacks against an exposed service or newly allowed high-risk traffic, however, deserve a specific investigation.
For a safe assessment of sources and countries, see Blocking malicious IP addresses and countries. If a packet was dropped, Analyzing dropped packets on Sophos Firewall leads from the Log Viewer and Rule ID to the actual cause of the drop.
Assess failed admin sign-ins
Review failed administrative sign-ins by time, source IP, destination service, username, and repetition. A typing error from the management network requires different treatment from distributed attempts from the internet or repeated sign-ins to a disabled account.
Open the Log Viewer from the upper-right corner of any web admin page; it opens in a new full-screen window. Select the appropriate module, set the period with Timer filter, and use Add filter to specify a field, condition, and value. Free-text search is useful for an IP address, username, port, or rule. Before making further changes, preserve the filtered entries as CSV with Export; Reset then clears all filters. A missing session entry doesn’t always prove that there was no traffic because firewall rules normally log sessions only when the firewall receives the connection Destroy event.
For suspicious attempts, check exposure and identity first:
- Is WebAdmin, SSH, User Portal, or VPN Portal intended to be reachable from the affected zone?
- Does the source come from an approved management network or a targeted Local Service ACL Exception?
- Is MFA active for the affected administrative access path?
- Do CAPTCHA, session timeout, and Block login work as planned?
- Are there simultaneous configuration changes or successful sign-ins for the same account?
Check network access under Sophos Firewall Device Access and Local Service ACL. For accounts, profiles, and offboarding, see Local administrators and device access profiles, and for the second factor, Enabling MFA on Sophos Firewall.
⚠️ Block login can block the source IP for multiple services after failed attempts. Don’t tighten values aggressively during an incident while there is no tested alternative admin and recovery path.
Understand HA, reporting, and model limits
In an HA cluster, each node stores only the logs and reports for the traffic it processed itself. For an event, therefore identify the node that was active or processing traffic at that time. An empty local report on one node doesn’t prove that no event occurred in the cluster.
Sophos Central Firewall Reporting can provide a consolidated view and longer retention. However, the local and central views aren’t treated as identical real-time sources. Enabling and operating Central Firewall Reporting explains log selection, arrival, and retention.
Additional limits:
- Control Center reports are updated periodically and aren’t a real-time event display.
- After upgrading from SFOS 20.0 or earlier to SFOS 21.0 or later, the Reports widget may show zero or a lower number until the next 24-hour update because Sophos stores reports from before and after the upgrade in separate databases.
- XGS 87/87w and XGS 88/88w don’t support on-appliance reports. Central logs, SIEM, and monitoring are therefore more important on these models.
- Missing data can be caused by logging, report period, licensing, retention, disk watermark, or the wrong HA node. It doesn’t automatically prove that no traffic was present.
Compare the firmware build with known issues
When the observation doesn’t match the expected state, compare the build recorded in the Control Center with the SFOS 22.0 release notes and the official Known Issues List. The exact build and symptom matter, not just the major version. Examples for the daily check include:
- NC-181971: On SFOS 22.0 GA and later, the IPS service may, in rare circumstances, enter a Dead state and fail to restart. Sophos doesn’t publish a self-service workaround and instructs customers to contact Support for the workaround.
- NC-181748: On SFOS 22.0 GA Build 411, Web Instant Alert emails aren’t generated for categories blocked by web policies. On this build, a missing alert therefore isn’t evidence that no block occurred.
- NC-180066, NC-180110, NC-178745, and NC-172912: The release notes list fixes in SFOS 22.0 MR2 Build 546 for stopped antivirus services, failsafe mode caused by the logging daemon, HA restarts caused by an out-of-memory condition, and flickering system graphs. If the symptom matches on an earlier build, record the issue ID and upgrade path in the ticket. Only update firmware in an approved maintenance window with a backup and rollback path.
Known issues can change independently of this article. Reopen the entry and preserve its current status before escalation. A matching issue ID can explain a symptom, but it doesn’t replace an impact assessment or functional verification.
Document and escalate deviations
A daily check is only complete when relevant deviations have a next step. For a ticket or operations journal, these fields are usually enough:
- date, time, and time zone;
- firewall name, model, SFOS version, and build;
- for HA: node, role, and last status change;
- affected function, zone, interface, VPN, or rule;
- observed state and expected state;
- screenshot, report period, log filter, or Rule ID;
- impact on users or services;
- owner, priority, next check, and escalation path.
Before any state-changing action, preserve the detailed status from the Control Center, the graph with its visible time range, and the filtered log lines or CSV export. For a support case, go to Diagnostics > Tools > Consolidated troubleshooting report to create a CTR with a System snapshot and the required log files: enter the reason, select Generate, and download the encrypted file when complete. Debug mode and log purging aren’t part of the daily check. They alter diagnostic state or destroy evidence and should only be used deliberately under Support guidance.
Immediate escalation is appropriate when a production WAN link or critical VPN path fails unexpectedly, a protection service is stopped, disk usage continues to grow, load remains high, repeated admin attacks correlate with a successful sign-in, or a new security event matches allowed malicious traffic.
Observation is more appropriate for a brief explainable spike, an intentionally unused interface, or a known event that already has a documented owner and stable functional verification.
Choose a useful check frequency
Sophos doesn’t prescribe one universal daily schedule for every view. The frequency therefore follows risk, operating hours, and existing monitoring:
- Daily or per shift: new messages, stopped services, WAN/VPN/HA, uptime, critical security events, and failed admin sign-ins.
- Weekly: graph trends, interface errors, disk growth, report patterns, recurring sources, and open tickets.
- After changes, upgrades, or failover: retest the affected function, logs, reports, alert path, and actual traffic.
- Regularly outside the short check: Health Check, firewall rule review, backup restore test, license and certificate expiry, and capacity planning.
Email notifications or monitoring shorten response time but don’t replace the review. An alert path is reliable only after transport, event selection, recipient, and response have been tested. The complete workflow is available in Setting up and testing Sophos Firewall email notifications.
Operations checklist
- Check the Control Center for new messages and unexpected status changes.
- Compare services, production WAN links, interfaces, VPNs, and uptime with the expected state.
- Read CPU, memory, load average, disk, and important interface counters against the baseline.
- Check security reports for the latest fully available period.
- Set the report period, select Generate, and export noteworthy results when needed.
- In the Log Viewer, select the module, Timer filter, and Add filter; correlate source, destination, user, action, and Rule ID, and preserve relevant lines as CSV.
- Assess failed admin sign-ins by source, service, and repetition.
- For HA, take account of the processing node and node-local logs.
- Compare the build and matching symptom with current release notes and known issues; record the issue ID in the ticket.
- Don’t trigger restarts, log deletion, or broad rule changes without preserving evidence and a recovery path.
- Document every relevant deviation with owner, priority, and next step.
- After a correction, retest not only the status but also the actual function.