Diagnose the Sophos NDR Integration Appliance and Sensor
This runbook isolates faults affecting the NDR Integration Appliance and NDR sensor. It begins with the exact status in Sophos Fusion, explains the visible status and metrics, and identifies when to collect logs for Sophos Support.
A status Connected or green is only an intermediate check. It proves neither full mirror coverage nor a successful upload of each dataset or a working end-to-end detection.
Quick procedure
- Record the appliance name, System ID, error start time with time zone, and exact message text.
- In Sophos Fusion, check the colour and appliance status, but do not restart anything yet.
- Review Status, NDR, Integrations, and Advanced in Appliance Manager and take timestamped screenshots.
- Assign the symptom to a fault class: platform/CPU, egress/upload, SPAN, registration, or shared resources.
- Only perform a reversible correction within this error class.
- Re-check the same measuring points under comparable load.
- If signals conflict, containers are not ready, or the action has no effect, collect logs and escalate.
Symptom, check, and next step
| Visible symptom | Record first | Check | Do not do |
|---|---|---|---|
Red: NDR containers not ready, <specific container names>. | named containers, Advanced, CPU platform, version, uptime | check CPU requirements and visible container status; then collect diagnostic data for support | do not modify containers manually or run kubectl or Dragonfly commands |
Red: Upload to s3 failed. Request was received but an error code was returned. Error code: <S3 upload error> | full error code, NDR upload, proxy/firewall changes | check DNS, routing, TCP 443, Web Proxy, and current Sophos egress destinations | do not create broad internet access or invent suspected individual hosts |
Red: spanX: unhealthy span | affected port, flow history, last mirror change | Check source, direction, destination, cable/port group, VLAN and tunnel path against the released plan | no CPU increase as a replacement for a wrong mirror configuration |
Yellow: spanX: packets being dropped | Port, time, CPU per core, traffic profile, further integrations | Check capacity and duplicate mirror sources; the message means more than 10% of packets are being dropped | Do not treat the threshold as an acceptable loss budget |
| Green, but no expected data or detections | capture, flows, upload, and tested path separately | check coverage, VLAN tags, source/direction, and the end-to-end test in stages | do not treat green or at least 2% unicast as proof of coverage |
| Connected, but no data in the Data Lake | Check NDR/integration upload and Advanced | Check egress and visible Dragonfly state; align with Pending CPU/EVC | No direct Dragonfly query or database change |
| Appliance remains Waiting for deployment | VM boot, MGMT address, DNS/NTP, Egress and correct appliance mapping | Check management path and bootstrap; assign image/seed only to the created appliance | No second manual registration and no unchecked rebuild |
| Appliance Manager is not reachable | Fusion status, MGMT IP, route, and applicable access rule | separate the management path from the SPAN path, verify the destination address, and follow the credential branch below | do not improvise a management IP on the SPAN interface |
| High CPU without further warning | per-core values, packet drops, upload, flows | distinguish expected DPDK cores from additional load | not automatically consider a single core as a fault at 100% |
Interpret red, yellow, and green correctly
Red: Integration does not work
For a red status, the specific message is more important than the colour:
NDR containers not ready, <specific container names>.means that at least one required application is not ready. Ifdragonflyis named or is conspicuous in the visible state, the check for CPU compatibility and platform requirements begins.Upload to s3 failed. Request was received but an error code was returned. Error code: <S3 upload error>means that the appliance wanted to upload to an S3 bucket via a pre-signed URL and received an error code. This is primarily an egress/proxy error class.spanX: unhealthy spanassigns the fault to the input of the named SPAN port. Check the sending network component or virtual mirror path first.
Yellow: Integration works with errors
spanX: packets being dropped appears when more than 10% of network packets are dropped. Packet capture and processing are CPU-intensive. A VM may require additional vCPUs; on certified hardware, the load can be distributed to another appliance according to the approved coverage plan. Co-hosted Log Collectors can place additional demand on the same resources.
A capacity change alone does not fix overlapping mirror sources, an oversubscribed destination path, or an incorrect SPAN configuration. Compare the traffic profile and topology before scaling.
Green: No currently reported integration error
Green means that NDR is receiving SPAN traffic and processing packet data without a reported problem. The current health classifier requires at least 2% unicast packets for a healthy SPAN port. This does not prove that every required VLAN, site, direction, or time window is captured. The previous statement that a port must show 100% unicast is not used.
For the detailed interpretation of health and capacity signals, see “Sophos NDR Health and Capacity Monitor”.
If expected detections remain absent despite a green status, first run the safe NDR test detection. If that check also fails and you suspect a VLAN issue, preserve the time window, SPAN port, selected VLAN, and visible appliance status for Sophos Support. Change VLAN Strip only after analysis confirms that both the selected VLAN and VLAN0 reach the sensor; otherwise leave the setting unchanged.
Narrow down local events with NDR Query
When capture and upload need to be checked separately, NDR Query in Appliance Manager can show the local event layer. The query runs against the NDR event database on this appliance VM, not against the Sophos Data Lake. Keep it clearly separate from the independently deployed Investigation Console, which provides data from an assigned NDR appliance for threat hunting on the local network.
- Record the affected appliance and incident time window.
- Open NDR Query and select Example queries on Query.
- Use Copy to copy a suitable predefined query, paste it into the text box, and run it with Go.
- Save the result under Query Results with a timestamp; reorder columns by drag-and-drop if needed.
- Compare the result with activity under NDR and the Fusion status for the same time window.
Appliance Manager currently supports only predefined queries here. Do not use a custom SQL query or a query from the Investigation Console. A local result when data is absent from Fusion directs further checks toward upload and egress. An empty local result directs them first toward the SPAN input, time window, and choice of predefined query; on its own, it is not proof of a fault.
Assess unexpected Nmap detections
If new Nmap-based OS scans appear in other security products, check the state of OS Detection under Global NDR Settings. When enabled, this option, which is off by default, scans every internal IP address seen by NDR every two hours. This can trigger detections in other security products.
Compare the activation time, affected appliance, target IP addresses, and detection timestamps. If activation was not approved, targets are unauthorized, or there is operational impact, turn off OS Detection again and document the time. Then monitor for any new scan events caused by this feature; process existing alerts according to the procedure for the respective tool. Do not repeat Nmap commands manually or create exceptions in other security products merely to conceal the symptom. Keeping the feature enabled requires documented approval from the network and security owners and validation over at least one complete two-hour interval.
Container and Dragonfly symptoms without CLI
Dragonfly processes NDR data. Two visible patterns are relevant for diagnosis:
- A red message with non-operational containers may occur if
dragonflyhangs in a restart loop because required CPU instructions are missing. - If the appliance is in Fusion Connected, but data does not reach the Data Lake and Dragonfly is on Advanced on
Pending, the EVC mode must be checked for a VMware EVC cluster. Sophos requires Skylake generation or later; Sandy Bridge is not supported.
For NDR VMs on VMware ESXi or Hyper-V, the pdpe1gb and avx2 CPU flags must be available. pdpe1gb is required for packet capture, avx2 for machine learning functions. More vCPUs do not compensate for missing flags. Hyper-V does not support Processor Compatibility Mode. For ESXi, VM Hardware version 11 or later and the documented platform requirements apply.
Safe testing:
- Document under Advanced status and visible name of the affected container.
- Under Status, record CPU usage, memory, root disk, and data disk.
- Compare the hypervisor, CPU model, EVC or compatibility setting, and flags exposed to the VM with the approved platform documentation.
- Correct an incorrect hypervisor/CPU setting only in a scheduled maintenance window; record initial value and return path beforehand.
- Then validate VM and appliance state over the normal operating surfaces.
- If
dragonflyremainsPending, a container is not ready or a restart loop is visible, create log bundles and escalate.
The platform values and supported CPU requirements are summarized in “Select and Size the Sophos NDR Platform”.
Check S3 upload and outbound connectivity
An S3 upload error does not mean that no SPAN traffic arrives. Capture and Upload are two separate stages. In Appliance Manager under NDR, therefore, record capture/flow activity and Uploaded for the same time window.
Check the Egress path in this order:
- Does the MGMT-IP configuration – DHCP or manual – match the management network?
- Does DNS resolution and NTP work over the intended services?
- Does the default route run on the intended Internet or central Egress path?
- Do Network ACL, Security Group or local firewall allow outbound HTTPS traffic?
- Does the Web Proxy allow the appliance and the required destinations without modifying or blocking the pre-signed S3 request?
- Do the rules match Sophos’s current port and domain exceptions?
Do not copy the region-dependent list of non-wildcard domains from old tickets. Compare it with the current Appliance requirements at the time of testing. Temporary broad access to the entire internet is not a safe test. Changes are made individually and validated with the same error time window after each step.
Rollback: A test-adjusted proxy, ACL or firewall rule is reset to baseline after the test, unless it is permanently required. When resetting, the previously functioning management and upload paths of other integrations must remain.
SPAN unhealthy, packet drops or missing flows
spanX: unhealthy span
Check for exactly the port mentioned:
- expected mirror source and direction,
- dedicated target interface and physical cabling,
- assignment of capture NIC, port group or vSwitch,
- for ERSPAN destination address, routing, MTU and GRE or VXLAN values,
- recent changes to VLAN, trunk, port group, host placement or mirror session,
- whether the same target is inadvertently mirrored again as a source.
Create harmless unicast traffic with a pilot host and observe the planned SPAN port and flow course in the same time window. If the activity is missing, the diagnosis remains with source, direction, filter, transport or capture allocation. Assess processing and upload only after traffic arrival at the input has been demonstrated.
spanX: packets being dropped
For drops above 10%, also record:
- allocated vCPUs and CPUs per core;
- bandwidth, packets/s, and flows/s,
- newly added or overlapping mirror sources,
- Memory and root and data disk trends,
- All log collectors of the same appliance with Received, Filtered, Accepted and Uploaded.
For a shared appliance, sizing starts with NDR and then accounts for the collector load. Additional appliance-wide limits are 8,000 collector events per second and, with 16 GB RAM, no more than 2 GB for Log Collectors. Other integrations can share even the CPU cores used by NDR and thereby affect NDR capacity. If collector load must be distributed, first define ownership and the destination appliance and use the generic integration guide; this runbook does not change vendor-specific syslog sources.
For 4 vCPUs, DPDK typically keeps one core at 100%, for 8 vCPUs, two cores remain at 100%. That alone is normal. A capacity problem is demonstrated by the combination of drops, additional saturated cores, declining upload, or a changed flow profile.
The Mirror chain, pilot validation and a limited return route are described under Plan and Validate Traffic Mirroring for Sophos NDR.
Diagnose registration and Connected
A new appliance initially appears as Waiting for deployment. After successful bootstrap and management path, the status of this appliance changes to Threat Analysis Center > Integrations > Configured > Integration Appliances under Connected.
If the status does not change:
- Identify the right appliance by name, platform and generated image or seed,
- check the VM boot process for persistent errors or restart loops,
- check the MGMT IP, VLAN, DHCP or manual values, gateway, and DNS,
- Check NTP and required egress against current appliance requirements,
- For ESXi, use the fusion-generated OVA only for a deployment attempt. For other platforms, check the current deployment process.
- Document time, visible state and last bootstrap output without secrets.
Do not delete or reinstall the appliance. This runbook deliberately does not contain a teardown or replacement process. Connected confirms the central connection and assignment, not SPAN coverage, upload or detection.
If the appliance was already Connected and loses this state, management path, egress and appliance availability are first checked. Mirror settings are not the first approach to this, because SPAN and management are separate paths.
Check Appliance Manager access instead of sensor errors
If Open Appliance Manager opens but sign-in with zadmin fails, treat this first as a sign-in problem, not as evidence of a SPAN, upload, or Dragonfly fault. In the Open Appliance Manager confirmation dialog, use reset it to set a new password. Store the new password directly in the password management system; neither the old nor the new password belongs in a screenshot, operations log, or support case.
If too many incorrect password attempts have locked the account, the documented alternative is the web console of the hypervisor hosting the appliance: select Unlock Account in the Weblink interface there. This fallback requires pre-existing authorized access to that hypervisor web console; this runbook adds no shell, SSH, or console commands and does not infer another access method from it. Then test sign-in once using the securely stored password. If the account remains locked, do not try more passwords; record the time and visible message and contact Sophos Support.
Offline management configuration as a final local recovery step
Change Actions > Settings > Management locally in Appliance Manager only when the VM has no network connectivity. If connectivity exists, make the change in Sophos Fusion. The VM being offline does not itself create a new access method: a local correction requires an existing, approved recovery procedure. Without such access, capture the current MGMT IP, last known Fusion status, and platform details, then escalate.
Before Save, compare the old and new values for IP Assignment, IPv4/Netmask, Gateway IP, DNS, DNS 2, and, where applicable, Enable Web Proxy, Web Proxy Type, Proxy URL, and Port Number. Proxy credentials remain in the password management system. Change only the value proven to be incorrect. If the interface requests confirmation for a restart, this is an appliance restart: NDR and all log collectors are interrupted. Therefore, first document shared workloads, the maintenance window, expected new IP address, and rollback path.
After the change takes effect, check reachability at the new IP address, the Fusion status, NDR capture and upload, and all log collectors. If approved recovery access remains available and validation fails, revert exactly the last change to the recorded baseline values. If the interface is no longer reachable, do not guess addresses or proxy values; escalate with the baseline, time, and impact. The complete precheck and post-check are in “Operate the Sophos NDR Appliance and Sensor Safely”.
Protect other integrations on a shared appliance
Before any restart or resource change in Fusion, open the arrow next to the appliance name and capture all integrations operated on the same appliance. In the Appliance Manager, Integrations shows their status, last restart and syslog counter.
- A single log collector can be restarted independently via Restart; NDR and other integrations remain active.
- Restart All affects all log collectors, but not NDR.
- Restart NDR affects the NDR sensor, not the log collectors.
- Actions > Restart affects the entire VM and interrupts NDR as well as all log collectors.
- Actions > Shutdown stops the entire VM and all integrations; a separately tested power-on path is required.
A broad VM restart is not a first diagnostic step. First, capture the status and metrics and determine the smallest affected component. “Operate the Sophos NDR Appliance and Sensor Safely” describes the impact and safe operating sequence.
Collect diagnostic data and logs
Baseline package
Record before making a change:
- Appliance name, System ID, Version, K3S Helm Chart version and Uptime,
- Fusion state and exact error text,
- start, time of reproduction and time zone,
- Under Status CPU per core, memory, root disk and data disk
- under NDR Capture per configured SPAN port, Uploaded and flow history,
- under Integrations all log collectors operated on the same appliance together with status and counters,
- container status visible under Advanced and last visible restart time,
- Platform, VM resources, traffic profile and recent changes,
- expected result, actual result and business impact.
If no appliance manager access is possible
- Open in Fusion Threat Analysis Center > Integrations > Configured > Integration Appliances.
- Select Collect logs in the three-point menu of the affected appliance.
- Open the information in the column Log requested and record the file name displayed there.
- Pass this filename with appliance, time window and error text to Sophos Support.
When Appliance Manager is available
- Select Open Appliance Manager in the three-point menu and then Open.
- Select Actions > Download Log File in the Appliance Manager.
- Only transmit the log archive to the existing case via the agreed support channel.
Log archives may contain IP addresses, host names and other confidential operational data. zadmin password, tokens, private keys, proxy credentials and other secrets never belong in ticket, screenshot or attachment.
Enable Remote Assistance in a controlled manner
Activate remote assistance only for a specific support case. The appliance must be online.
- Open in Fusion Threat Analysis Center > Integrations > Configured > Integration Appliances.
- Select Remote Assistance in the three-point menu.
- Activate Enable in the dialog.
- Set the Sophos Group Privacy Notice Confirmation and select Save.
- Wait until a Access ID is displayed.
- Only send this Access ID to Sophos Support via the agreed channel.
Access ends automatically after no more than seven days. If the analysis is completed earlier, deactivate Enable in the same dialog and document the end. Remote Assistance does not replace a support case or diagnostic package.
Correction safely validate and roll back
Change only one hypothesis per run. Record the baseline, responsible person, maintenance window, and rollback path first. Then test under comparable load:
- previous red or yellow message does not occur again,
- the expected appliance state and management path are stable,
- each intended SPAN port shows activity matching the pilot traffic,
- Flow history and upload remain stable over a meaningful time window,
- no message for more than 10% drops returns,
- CPU outside the expected DPDK cores, memory, and storage show sufficient headroom,
- all log collectors operated on the same appliance continue to process data,
- A platform change provides the required CPU flags and supported mode.
If verification fails or new effects arise, reverse only the most recent change. If the baseline cannot be restored, make no further changes, collect diagnostic data, and open a Case with Sophos Support.
After the technical correction, a green integration is still not proof of detection. Only after the mirror and upload chain is stable should you follow “Generate and Verify a Safe Sophos NDR Test Detection”.
Escalate to Sophos Support
Open a case with Sophos Support if:
NDR containers not readyremains ordragonflyremains visible inPendingor a restart loop,- required CPU flags are not available despite the correct platform,
- an S3 upload error remains despite confirmed DNS, proxy, firewall and egress path,
spanX: unhealthy spanremains despite verified source, direction and destination allocation,- Packet drops recur after suitable capacity or load distribution,
- Connected, local upload and data lake reception contradict each other,
- the bootstrap does not reach registration or the appliance unexpectedly switches between states,
- a secure correction would require low-level container, Kubernetes or Dragonfly interventions.
In the Case, send the baseline package, the log file name or the log archive, exact steps and measurable results. Clearly list hypotheses that have already been ruled out. Sophos Support handles product problems in installation, administration, and operation; the Case is not a request to investigate a Detection.
An XDR-based, Self-managed Detection remains the customer’s responsibility. Only a Sophos-managed MDR Case is investigated and handled by Sophos MDR. For Case creation and escalation, see “Open a Sophos Support Ticket with Support Assistant”.