Monitor Sophos NDR health and capacity
A green Sophos NDR status is an important interim check, but it is not proof of complete mirror coverage or a functioning detection chain. For a reliable assessment, six groups of signals must be considered separately: Sophos Fusion, Appliance Manager, SPAN input, upload, compute and storage, and—if present—the standalone Investigation Console.
Quick check: First, review the NDR status in Sophos Fusion. Then, in Appliance Manager under NDR, check the values for every expected SPAN port, Uploaded, and the flow history. Under Status, check CPU, Memory, Root Disk, and Data Disk. A yellow spanX: packets being dropped notice means that more than 10% of the network packets processed by NDR are being dropped. A green SPAN port currently meets the product classifier of at least 2% unicast packets. None of these statements by itself proves that all intended networks are being mirrored or that a detection is reaching the end of the chain.
Match the indicators to their purpose
| Signal | Location | What it demonstrates |
|---|---|---|
| Red, yellow, or green | Sophos Fusion, NDR integration | Aggregated integration status and specific status message |
| NDR | Appliance Manager | Upload percentage, captured percentage for each configured SPAN port, and Network Flows in 30-second intervals |
| Status | Appliance Manager | CPU, Memory, Root Disk, and Data Disk utilization of the appliance |
| Integrations | Appliance Manager | Status and syslog counters of third-party integrations running on the same appliance, not NDR SPAN traffic |
| Investigation Console | Separate component | Its status does not demonstrate the status of the NDR integration or the SPAN coverage |
Open Appliance Manager in Sophos Fusion by going to Threat Analysis Center > Integrations > Configured > Integration Appliances. On the appliance row, select the three dots and then Open Appliance Manager. The header area shows, among other details, Version, K3S Helm Chart version, Uptime, and System ID. Include these details in every incident record because two appliances with apparently similar symptoms may be running different software versions or may have different uptimes.
Use Fusion status as the starting point
Sophos Fusion distinguishes between three states:
- Red: The NDR integration is not working.
NDR containers not ready, <specific container names>indicates applications that are not ready.Upload to s3 failed ...relates to the cloud upload.spanX: unhealthy spanrelates to the traffic input from the mirroring network device. - Yellow: The integration is working, but with errors.
spanX: packets being droppedis displayed when more than 10% of network packets are being dropped. - Green: The integration is receiving SPAN traffic and processing packet data without a reported error. A SPAN port is currently classified as healthy if at least 2% of the observed packets are unicast packets.
The two percentages have different denominators and must not be offset against each other. The 2% threshold classifies the composition of the traffic arriving at the sensor. It does not mean that 2% of all corporate traffic is sufficient, nor that NDR may lose 98%. “At least 2%” includes exactly 2%. By contrast, the drop message describes the proportion of packets already arriving at NDR that cannot be handled by processing because resources are insufficient. Sophos documents the warning for more than 10%, not for “10% or more.”
Important: The threshold of more than 10% is a product status, not a target or an acceptable loss budget. Even a value below the reporting threshold can represent deterioration compared with your own baseline. Likewise, a green port can deliver incorrectly or incompletely mirrored traffic as long as the observed mix satisfies the unicast classifier.
Check SPAN input and upload separately
In Appliance Manager, the NDR tab shows a capture value for every configured SPAN port. SPAN Port 2 appears only if a second port has been configured. The overall Network Flows graph is displayed in 30-second intervals.
These indicators answer three separate questions:
- Is traffic arriving on every expected SPAN port? A missing or suddenly and significantly different value directs the investigation first to the source switch, the mirror or SPAN session, the virtual network interface assignment, and any recently changed VLANs or port groups.
- Is the traffic mix plausible? Green confirms only the current unicast classifier. The flow history must also be consistent with usual times of day, locations, and expected traffic peaks.
- Can NDR process the packets? The yellow drop message indicates a processing bottleneck. It is not the same as an oversubscribed mirror port or packet loss in the production path.
The upload percentage, also visible on the NDR tab, represents a downstream stage. A healthy SPAN input with a declining upload does not indicate the same class of failure as an empty SPAN port. For Upload to s3 failed. Request was received but an error code was returned, check the appliance’s outbound internet access as well as firewall and web proxy rules. NDR uploads the data to an S3 bucket using a presigned URL. If the error persists after correcting the network or proxy configuration, contact Sophos Support.
For third-party log collector integrations, the cards under Integrations count Received, Filtered, Accepted, and Uploaded. These syslog counters are neither the NDR upload value nor the SPAN capture value. They are still important because a heavily loaded log collector integration running on the same appliance consumes that appliance’s CPU and Memory.
Assess CPU, Memory, and storage
Under Status, Appliance Manager shows CPU, Memory, Root Disk, and Data Disk utilization. Individual CPU cores being continuously fully utilized is expected with NDR: the Data Plane Development Kit (DPDK) operates on reserved cores in poll mode. It continuously polls for packets instead of waiting for interrupts while idle.
Sophos gives two specific examples:
- On a VM with 4 CPU cores, one core remains at 100%.
- On a VM with 8 CPU cores, two cores remain at 100%.
This type of per-core utilization is therefore not, by itself, evidence of overload and does not necessarily disappear when traffic is low. Conversely, “DPDK is normal” must not be used to explain every instance of high CPU utilization. A combination of increasing traffic, additional fully utilized cores, the packets being dropped message, declining upload, or a flow history that has changed from the baseline becomes critical.
The following capacity guidelines apply to virtual appliances:
| Traffic profile | Documented upper limit | Sizing |
|---|---|---|
| Medium | up to 500 Mbit/s, 70'000 packets/s, and 1'200 flows/s | The VM’s default settings can be used |
| High | up to 1 Gbit/s, 300'000 packets/s, and 4'500 flows/s | Scale the VM to 8 vCPUs |
All three metrics must be considered together. An environment can remain below the bandwidth limit but still generate many packets per second because the packets are very small. If the values exceed the High profile, Sophos recommends deploying multiple virtual appliances in the network.
These values apply to a VM running only NDR. Under high load, each additional log collector integration hosted on the appliance requires approximately 400 MB of RAM and may consume additional CPU capacity. For this reason, also check the cards under Integrations before scaling up. With a sustained mixed workload, separating the workload across multiple appliances may be more appropriate than repeatedly adding resources to the same VM.
The product documentation used for this article does not specify a general warning threshold for Memory, Root Disk, or Data Disk. It would therefore be misleading to treat an arbitrary percentage as a Sophos limit. The relevant factors are the trend, available headroom, and concurrent symptoms. If disk usage grows continuously, do not manually delete files or containers. First document the status and time period, and if the cause is unclear, preserve logs for Sophos Support.
Establish an actionable baseline
A snapshot cannot distinguish a normal daily pattern from the onset of overload. After deployment and after every relevant change, therefore, record comparable data points:
- date, time, and time zone, as well as the expected load period,
- Fusion color and exact message text,
- capture value for every configured SPAN port,
- upload percentage and the shape of the flow history,
- CPU per core and overall, Memory, Root Disk, and Data Disk,
- vCPUs and RAM assigned to the VM,
- estimated Mbit/s, packets/s, and flows/s, or values measured on the source system,
- third-party integrations running at the same time and their activity,
- changes to the switch, hypervisor, proxy, firewall, or appliance.
Measurements during a quiet period, normal business load, and a known peak are useful. The goal is not a universal target value, but comparison of the same appliance under similar conditions. This makes a sudden decline on a SPAN port visible even if Fusion is still green. After a capacity change, establish a new baseline only once the state is stable.
Resolve issues in a safe order
- Record the scope: Document the affected appliance, SPAN port, start time, exact Fusion text, and most recent change. Preserve screenshots or measurements before a restart.
- Check the input: For
unhealthy span, missing flows, or a deviation from the port baseline, first check the mirror source, destination port or virtual network interface, and the expected networks. Additional CPU does not fix an incorrectly configured SPAN source. - Check the upload: For an S3 upload error, check outbound internet access, the firewall, and the web proxy. Conversely, a successful upload does not fix missing mirror coverage.
- Check capacity: When drops exceed 10%, compare the traffic profile, other fully utilized cores, vCPU allocation, and integrations running on the same appliance. For a VM, assign additional vCPUs; the documented High profile specifies 8 vCPUs. For certified hardware, another appliance can be deployed and the SPAN traffic split between them. Before splitting the traffic, an approved coverage plan is mandatory: it must clearly assign every intended network, VLAN, and mirror source to the target appliance and prevent gaps as well as unintended duplicate feeds. For mixed workloads, NDR and log collectors can be distributed across separate appliances.
- Define the scope of a restart: If NDR still does not operate correctly after the cause has been resolved, a targeted NDR restart can be considered in accordance with the applicable operations or troubleshooting instructions. Normal NDR processing does not take place during the restart; a restart is not a substitute for correcting capacity. Restarting or shutting down the entire VM has a greater scope of impact and does not belong in the first troubleshooting step.
If vCPUs or the traffic distribution are changed, do so in an approved maintenance window and in accordance with the requirements of the virtualization platform in use. Make one change at a time so that its effect remains measurable. After splitting traffic, check the coverage plan against every resulting SPAN port; also validate the approved end-to-end test path.
Do not confuse the Investigation Console with the NDR appliance
The Investigation Console is a separate component. Its status demonstrates neither the status of the NDR integration nor the completeness of the SPAN sources or successful data delivery. This article is therefore limited to clarifying the boundary: SPAN, NDR upload, DPDK, and packet drops are checked at the NDR integration and in Appliance Manager; checking and troubleshooting the Investigation Console belongs in the applicable console operations documentation.
Validate after every action
Repeat the checks under a load comparable to the initial measurement. A correction is considered effective only when:
- Fusion shows the expected state without the previous red or yellow message,
- every expected SPAN port is visible and once again matches its own baseline history,
- the unicast classifier has not been incorrectly used as proof of coverage,
- the message for more than 10% packet drops does not recur,
- upload and Network Flows remain stable over a meaningful observation period,
- CPU outside the expected DPDK cores, Memory, Root Disk, and Data Disk show sufficient headroom,
- third-party integrations running on the same appliance continue to process the expected data,
- after traffic has been split, every intended network, VLAN, and mirror source is fed to exactly the intended appliance in accordance with the approved coverage plan, and the approved end-to-end test path works.
These checks validate the operational state, but not yet a complete detection chain. End-to-end proof additionally requires an approved NDR test and verification of the resulting detection.
When to involve Sophos Support
Involve Sophos Support if containers do not become ready, an S3 upload error persists despite a confirmed internet and proxy path, a SPAN port remains unhealthy despite corrected source configuration, packet drops recur after an appropriate capacity increase, or resource indicators and status messages contradict each other.
Prepare at least the following data for escalation:
- appliance name, System ID, Version, K3S Helm Chart version, and Uptime,
- exact status and error text, including the start time and time zone,
- affected SPAN port, capture and upload values, and flow history,
- CPU per core, Memory, Root Disk, and Data Disk before and after the action,
- VM allocation, observed traffic profile, and integrations running on the same appliance,
- recent switch, hypervisor, firewall, or proxy changes,
- actions performed and their measurable results,
- for issues with a separate Investigation Console, the evidence required by its applicable operations documentation.
Passwords, private keys, and other credentials do not belong in the ticket. Do not run low-level Kubernetes commands or manually modify containers based on suspicion; for NDR containers not ready, collect the available observations and coordinate with Sophos Support.