Skip to content
Avanet

Planning Sophos Switch as the Access Layer for XGS HA

Sophos switches can form the access layer in front of a Sophos Firewall HA pair. The safe principle is simple: The Dedicated HA Link connects the two XGS firewalls directly and remains separate from all production switch paths. LAN, DMZ and, where applicable, WAN are distributed across two switches so that the failure of one switch does not disconnect both firewall nodes from the same network at once.

This runbook covers the switch and cabling side. Set up Sophos Firewall High Availability explains the choice between Active-Passive and Active-Active and the complete firewall configuration.

Important: Two firewalls alone do not eliminate a common switch, a common power supply, or a common provider path as a cause of failure. An HA status Active-Passive or Active-Active only confirms the firewall cluster, not the redundancy of the entire access layer.

Target architecture and safety boundaries

A robust design separates three types of traffic:

  1. Dedicated HA Link: direct physical connection between Primary and Auxiliary. Heartbeat and the synchronisation of configuration, status and sessions run on top of this. In Active-Active, the link is also used for internal HA traffic distribution between Primary and Auxiliary. However, it is not an ordinary LAN, DMZ or WAN interface and is not integrated into any of these production networks.
  2. Production paths: Connections from each firewall node to LAN, DMZ and WAN. Both nodes must be able to reach the same production networks after a role change.
  3. Management paths: Access to firewalls and switches for acceptance, troubleshooting and rollback. At least one independent access must not depend on the exact link that is currently being changed or tested.

The reference documented by Sophos uses two CS210-48FP switches for LAN and DMZ. Each firewall node connects to one of the switches; a direct interlink carries the LAN and DMZ VLANs Tagged between them. A separate CS110-24FP connects the WAN side of both firewalls and the internet handoff in an Untagged VLAN. The firewalls’ Dedicated HA Link is also connected directly in this example.

This example is not a general port plan. Its port numbers and VLAN IDs are only for understanding the roles:

PurposeReference valueMeaning
LANVLAN 100Firewall, LAN and switch interlink ports
DMZVLAN 200Firewall, DMZ and switch interlink ports
WANVLAN 300Both firewall WAN ports and Internet handover
Switch interlinkPort 52VLAN 100 and 200 Tagged
Dedicated HA LinkFirewall port 7Direct connection, not via the switches

Take your VLAN IDs, ports and interface names from the existing network plan. Do not copy the reference values into a production network where they already have a different meaning.

What the reference design does not make redundant

The reference single WAN switch remains a shared failure domain. If it fails, both firewall nodes lose this WAN path. If the WAN switching layer is also to survive a single device failure, it needs a separately tested dual switch and provider design; XGS HA does not create this redundancy by itself.

Even two switches mounted next to each other are not separate failure domains. To ensure a reliable separation, check at least:

  • separate power supply or separate fused current paths;
  • separate switch hardware and, where possible, separate rack or patch paths;
  • one firewall node per switch instead of both nodes on the same access switch;
  • separate cable routes without shared patch or transceiver risk;
  • reachability of the production networks through each switch individually;
  • monitoring for both switches and both firewall paths;
  • documented responsibility for switch, firewall and provider faults.

This assignment is not only recorded in the network plan, but is also continued as part of the design matrix right through to testing and rollback. This means it remains visible which power, rack, patch, switch and provider dependencies a specific path actually has.

VLAN, LAG and STP design together

The Layer 2 topology is fully defined before cabling. For each network, the plan includes VLAN ID, Tagged/Untagged role, PVID, involved firewall ports, switch ports, and the expected path after a failure. Tagged, Untagged and PVID are implemented in Configure Sophos Switch VLANs securely.

Keep VLANs consistent across both switches

In the Sophos reference setup, the firewall and network ports for LAN and DMZ are members of their respective VLANs. The interlink between the two CS210 switches carries VLAN 100 and 200 Tagged. The PVID of the Untagged ports corresponds to the respective VLAN; Sophos sets Ingress filtering: On and Accept type: All in the example.

The following test rules apply to your own design:

  • A VLAN must have the same VLAN ID at every link end involved.
  • Only the VLANs Tagged that actually have to reach both switches are permitted on the interlink.
  • Untagged VLAN and PVID of an access port must match.
  • The Dedicated HA Link is not included in any of these production VLANs.
  • Management VLAN and the rollback path remain accessible during the changeover.
  • Ingress filtering will only be tightened after VLAN membership and PVID are proven.

The reference is port-based and Untagged on the firewall interfaces. If your own firewall uses VLAN subinterfaces or trunks instead, you must not adopt the Untagged/PVID values from the example. Then the Tagged path on the firewall, switch and interlink must consistently match your own interface design.

Do not confuse LAG with chassis redundancy

A LAG bundles several links to one logical peer. Depending on the traffic distribution, it increases capacity and can absorb the failure of a member. But it does not prove that a complete switch can fail.

The two Sophos shared sources do not document a multi-chassis LAG over two independent Sophos switches for this XGS-HA example. Therefore, a firewall LAG is not simply distributed to Switch A and Switch B with one member each. Such a structure is only permitted if the entire switch solution explicitly works as a supported logical LAG counterpart to the firewall and the specific design is documented and tested separately.

Without this proof, separate firewall interfaces to separate switches are the safe planning limit. LACP, static LAGs, member activation and decommissioning are covered in Securely configure Sophos Switch Ports, LAG and Spanning Tree.

Plan STP in front of a redundant Layer 2 path

As soon as more than one Layer 2 path can be created between two switches or via downstream LAN/DMZ infrastructure, the loop-free topology must be established before plugging in the additional cable. RSTP is suitable for a simple shared topology; MSTP only for deliberately consistent region and instance design.

Before the change, record the desired Root Bridge and bridge priorities, forwarding and expected blocking ports, and the behaviour when the interlink fails and returns in the design matrix. For MSTP, also record the VLAN assignment for each instance. Reserve Edge ports for genuine end devices, not firewalls, switches or unknown bridge connections. Define an independent management path in case convergence fails.

Loopback Detection can help in addition, but does not replace STP. A blocking STP port is not automatically a failure in a planned redundant path.

Prepare for change

Before the first change, create a design matrix for both firewall nodes and both switches. It is the binding working document for implementation, testing, acceptance and rollback. The following compact example shows the structure. XGS-A, SW-A, the port numbers and VLANs are realistic example values, not product specifications; take different values from your own network and patch plan. Replace every item in square brackets with a specific value for your environment before approval.

Design matrix CHG-[number] — example status before the change
HA: Active-Passive | XGS-A = Primary/Active | XGS-B = Auxiliary/Passive | synchronised
Dedicated HA: XGS-A Port7 <-> XGS-B Port7 | directly | not part of the failure test
Management/rollback path: [separate admin path] | Responsible: [Name]
STP: RSTP | Root: SW-A [Priority] | Secondary: SW-B [Priority]
Backups/previous state: [Switch/Firewall storage] | Patch schedule: [version] | Rollback approval: [Name]

P1 LAN: XGS-A Port1 <-> SW-A 1/0/47 | VLAN 100 Untagged/PVID 100
no LAG | RSTP Forwarding, not Edge | Monitored Port: yes | Power path A
P2 LAN: XGS-B Port1 <-> SW-B 1/0/47 | VLAN 100 Untagged/PVID 100
no LAG | RSTP Forwarding, not Edge | Monitored Port: yes | Power path B
P3 Interlink: SW-A 1/0/52 <-> SW-B 1/0/52 | VLAN 100,200 Tagged
PVID [according to local native VLAN concept] | no LAG | RSTP Forwarding
[Complete DMZ, WAN and LAG paths in the same pattern; for a LAG, identify the logical peer
and name all members. Explicitly mark shared WAN/provider failure domains.]

T1 monitored port failover: Prerequisite XGS-A Primary/Active, XGS-B Auxiliary/Passive,
both synchronised; disconnect exactly XGS-A Port1/P1.
Expectation/Path: XGS-A is no longer processing traffic; XGS-B becomes Active;
LAN traffic passes through P2 and SW-B.
Validation: status of both nodes, synchronisation after return, switch ports/VLAN path,
Gateway + [internal destination] + [external path] + [critical application], timestamp.
Abort: both nodes Active, loss of management or traffic failure > [approved duration].
Rollback: reconnect P1, wait for synchronisation, check data path again;
Only then restore preferred roles in a controlled manner, no uncontrolled failback.
T2 switch failure: [Isolate SW-A] | Expectation/Path: Service via XGS-B/SW-B and P2
Validation: [same technical tests] | Abort: [criterion]
Rollback: Create SW-A/links individually, wait for STP to be stable and HA to be synchronised.

P lines are copied for other networks, and a separate T line is copied for each approved failure scenario. This means that the physical remote station, zone, VLAN/PVID, Tagged/Untagged role, LAG and STP status, Monitored Port, HA role and current status are not in separate checklists. Expected failure path, technical validation, abort and return path remain directly linked to the tested interface.

Before the maintenance window, the switch and firewall configurations referenced in the header, the effective port, VLAN, LAG and STP pre-state as well as the cable and patch plan must actually be secured and accessible via the specified independent administration path. Simply entering a storage location is not enough.

A switch backup does not replace documentation of the effective Layer 2 status. The difference between Fusion and local backup is explained in Sophos Switch Backup and Restore.

Build access layer in a controlled manner

1. Prepare switches without parallel loop

  1. Process the P lines of the design matrix one after the other: configure VLANs, PVIDs and the planned switch interlink first and confirm the actual status directly in the same line.
  2. Before enabling a redundant Layer 2 path, enable STP and check Root Bridge and port roles.
  3. Leave additional cables disconnected or keep their ports disabled.
  4. Clarify configuration source and conflicts between Sophos Fusion and local switch configuration.
  5. Retest management access after each change.

In the local Switch web interface, the Sophos reference for VLANs uses:

Configure > VLAN settings > 802.1Q

PVID, Ingress filtering and Accept type are edited there under PVID and ingress filter. Further UI steps and their security limits remain in the linked VLAN runbook so that two different instructions are not maintained.

2. Activate exactly one production path per network

First, only the clearly planned individual path is activated for LAN, DMZ and WAN. Then you check:

  • Link condition and agreed speed;
  • effective VLAN membership and PVID;
  • Management access to switch and firewall;
  • Reachability of the intended gateway;
  • permitted test traffic through the currently active firewall.

Only when this state is stable will an additional switch interlink, LAG member or second network path be activated individually. After each cable, STP role, LAG membership and management access are checked again.

The Dedicated HA Link is wired directly between the same designated ports on both firewalls. It is not routed through the production switch interlink, a LAN, DMZ or WAN VLAN, or a shared access switch. This keeps faults in the production Layer 2 topology separate from the heartbeat, synchronisation and internal HA distribution path. In Active-Active, this direct link can also distribute traffic between nodes for processing; that does not make it an ordinary production network connection.

The Sophos reference sets up the firewall in the Interactive mode: first the Auxiliary, then the Primary. Both use the same dedicated HA link port and passphrase. On the Primary, cluster ID, peer link address, Monitored Ports, Peer Administration and Keepalive values are set. These firewall fields are not blindly adopted from the reference example; The complete process and requirements can be found in the linked Firewall HA article.

4. Select Monitored Ports by failure domain

A Monitored Port is intended to detect a relevant path failure. The role or status change that results depends on the HA mode and which node or path is affected. In the Sophos example, the Primary monitors its LAN and DMZ ports. Only permanently connected, critical interfaces are selected for your own environment.

Unused, only temporarily active or intentionally disconnected ports are not useful. You would trigger an unnecessary failover. The Dedicated HA Link and a Monitored Port remain different interfaces with different tasks.

Acceptance and failure tests

The tests take place in a maintenance window. Before each failure test, ongoing production changes are stopped, current roles are noted and a person responsible for the immediate return path is appointed. Only one failure domain is changed at a time.

Each monitored-port test follows an approved T line such as T1: reconfirm the HA mode, affected node, configured role, current Active/Passive status, exact interface, expected result and abort condition immediately before disconnection. The normal Active-Passive failover test disconnects a monitored interface on the currently active, traffic-processing node; the peer must take over the traffic. For Active-Active or a path on the Auxiliary node, approve a separate T line with the actual expected behaviour instead of assuming the role change from T1.

Verify the baseline state

Before a failover, the normal state must be clear:

  • Both switches are reachable and their expected ports are stable.
  • VLAN memberships, PVIDs and Tagged interlinks correspond to the design matrix.
  • LAGs only contain the planned members.
  • Root Bridge, STP port roles and port states are consistent with design.
  • Both firewalls show HA scheduled mode and a synchronised state.
  • Dedicated HA Link and all selected Monitored Ports are active.
  • Test clients reach gateway, explicitly allowed internal destinations and the intended external path.
  • Peer administration or independent management access works.

Maintenance and failure testing in a safe order

  1. Disconnect a LAG member in a controlled manner, if present. The logical path must work over the remaining member; then add the member again and check that it rejoins.
  2. Disconnect Switch Interlink in a controlled manner. Only do this if the expected data path is clearly described in the plan without it. Check roles, VLAN reachability and STP health, then restore the link and observe reconvergence.
  3. Disconnect the monitored firewall path specified in the design matrix. In Active-Passive, the exact monitored interface of the currently active traffic processing node is used and it is checked whether the peer takes over as documented. Active-Active and Auxiliary path tests only follow their separately documented expected state. If there is a discrepancy, abort the test and restore the path. Then check synchronisation, status of both nodes and the same test traffic; do not allow an automatic failback to proceed uncontrollably.
  4. Disconnect or completely isolate an access switch in a controlled manner. The other switch, the associated firewall node and the intended networks must deliver the service expected in the design.
  5. Restore the preferred normal state in a controlled manner. Switch, links and firewall roles may only be returned after stable synchronisation and a verified data path.

The Dedicated HA Link is not disconnected as a normal availability test in the running production network. Its failure can result in both firewalls no longer seeing the peer. Such a split-brain test requires a separate, explicitly approved procedure with isolated production interfaces. For normal acceptance it is sufficient to check the link status, the HA synchronisation and the documented procedure for its failure.

A test is not considered passed simply because a ping continues. After each step, roles, synchronisation, switch-port states, VLAN path and site-critical applications are also checked. Observations and timestamps go into the operational log; Guaranteed switching times cannot be claimed without measurement.

Troubleshoot by symptom

HA is green, but a network cannot be reached after the role change

  • Compare VLAN ID and Tagged/Untagged membership on both switches.
  • Check PVID of the affected Untagged port.
  • Check whether the VLAN is actually approved over the interlink Tagged.
  • Check firewall port and physical cabling against the design matrix.
  • For a LAG, determine whether all members lead to the correct logical counterpart.
  • Check the STP port role; an unexpectedly blocked or disabled path can prevent accessibility.

Failover occurs unexpectedly

  • Check on the firewall which Monitored Port triggered the condition.
  • Examine link flaps, transceivers, cables and speed/duplex adjustment on the associated switch port.
  • Ensure that no optional or intentionally unconnected port is being monitored.
  • Check LBD and STP status before reactivating a port.

Both firewalls no longer see the peer

The Dedicated HA Link is the first boundary to check. Do not move production cables at random or restart both nodes at the same time. Decide which node should continue processing traffic, disconnect or shut down the other node from the production network in a controlled manner, and then check the cables, ports and link state of the direct HA link. Reintegrate the second node fully only after peer contact is stable and the roles are unambiguous.

After activating the second path, packet loss or a loop occurs

  1. Disconnect the last activated additional link in a controlled manner.
  2. Create a unique individual path again.
  3. Check the STP Root Bridge, port roles and port states on both switches.
  4. For a LAG type and member ports, compare on both ends.
  5. Check VLAN interlink list and possible unintentional Untagged connection.
  6. Only reactivate exactly one additional link after correction.

Switch can no longer be administered

Use the prepared independent management access. Undo the last change to Management-VLAN, PVID, Uplink, LAG or STP based on the previous state. Do not restart both switches as a precaution: this could result in the data path that is still working being lost.

Rollback

A rollback restores the documented previous state and begins at the last redundancy added:

  1. Stop test traffic and save current HA, switch and STP state.
  2. Disconnect the last activated additional link, LAG member or interlink in a controlled manner until exactly one unambiguous production path remains.
  3. Only return firewall roles if the intended normal path is stable.
  4. Undo LAG, STP and VLAN changes on both link ends in reverse order.
  5. Restore the previous Tagged/Untagged memberships and PVIDs from the design matrix.
  6. Repeat management access, HA synchronisation and the same functional test as before the change.
  7. Leave the Dedicated HA Link directly connected as is, unless this exact link is repaired using your own approved procedure.

If the previous state itself only had a common switch or a single uplink, the rollback deliberately restores its lower availability. This is documented in the change record as a remaining risk and is not described as a fully redundant outcome.

Operational checklist

  • Each P line of the design matrix contains the confirmed actual state; Deviations have been resolved or accepted as a residual risk.
  • Each T line contains result, timestamp and verifier. The tests were run individually, not as a combined failure.
  • Dedicated HA Link, HA roles, synchronisation, STP state and management access are back to the approved normal state.
  • Switch and firewall configuration, updated patch plan and the record of service tests are saved at the specified storage location.
  • Rollback path, responsible parties and remaining common power, patch, rack, WAN or provider failure domains are recorded in the change record.