Skip to content
Avanet

Systematically troubleshoot Sophos Firewall authentication errors

An authentication error on Sophos Firewall can occur at very different points. The server may be reachable but not selected for the affected service. Sign-in may work while the main group, quota, or firewall rule prevents access. Or the firewall may not see the user at all and processes the traffic only by IP address.

The safe quick path deliberately separates these layers:

  1. Record the user, source IP, affected service, time, and expected result.
  2. Test reachability of the portal or authentication service from the actual source zone.
  3. Under Authentication > Services, verify that the correct server is selected for the affected service and is in the intended position.
  4. Run Test connection on the authentication server, but treat it only as proof of credentials and server connectivity.
  5. Generate a fresh sign-in and check the user, source IP, and Client Type under Current activities > Live users.
  6. Under Authentication > Users, check status, main group, other groups, user overrides, and User ID.
  7. Check access time, quotas, MFA, and sign-in restrictions only for the user actually affected or their effective group.
  8. Only after identity succeeds, test the expected firewall rule, Rule ID, web or VPN policy, NAT, and routing with real traffic.

⚠️ Don’t broadly change authentication servers, group order, quotas, or Device Access on suspicion. A service restart, Purge AD users, or an Any rule isn’t a normal first diagnostic step. First prove at which layer the workflow stops.

Five layers instead of one sign-in error

In practice, a simple model helps. User access doesn’t pass just one check but at least five separate layers:

  1. Reachability: The client and firewall reach the portal, proxy, or authentication server through the intended path.
  2. Service and method: The correct authentication server is selected under Authentication > Services for this exact service.
  3. Identity: The firewall recognizes the expected user and shows them with the correct source IP and client type.
  4. Authorization: User status, group, main group, access time, quota, MFA, and service-specific permissions match.
  5. Subsequent traffic: A firewall rule, web policy, or VPN policy allows the required destinations, and the return path works.

This separation prevents two common false conclusions. A green Test connection doesn’t prove that SSO sign-in works. Conversely, a user under Live users doesn’t yet prove that their traffic matches the intended rule or policy.

Define a reproducible test case

Before the first change, define one test case. Safe documentation values include:

  • User: auth-pilot@example.com
  • Client IP: 192.0.2.25
  • Source zone: LAN
  • Service: Firewall authentication, User portal, VPN portal, SSL VPN, or Web admin console
  • Test time: local firewall time with date and minute
  • Expected group: Internet-Standard
  • Expected rule: LAN-Users-to-WAN
  • Expected result: successful sign-in and HTTPS access to a defined test destination

example.com and 192.0.2.0/24 are documentation values. Replace the username, source IP, zone, group, rule, and destination with real pilot values. The pilot should use the same authentication source and permission group as the affected user but have no unnecessary administrative privileges.

If possible, also test a working comparison user from the same source. If both fail, the cause is more likely the service, server, or reachability. If only one user fails, the user object, groups, quotas, MFA, or sign-in restrictions are more likely.

Check the service, reachability, and server

Check Device Access only for the required service

Portals and local authentication services aren’t allowed by normal firewall rules. Under Administration > Device access, or with a targeted Local service ACL exception rule, the required service must be reachable from the intended source zone or known source networks.

Different entries may be relevant depending on the test, such as Captive portal, User portal, VPN portal, SSL VPN, AD SSO, or Client Authentication. Check only the service actually required. Broadly allowing LAN or WAN would hide the fault and may expose additional management or portal services. Device Access and Local Service ACL on Sophos Firewall explains the safe boundary.

Reachability alone isn’t authentication. A visible portal page only confirms that the client reaches the local service. It doesn’t yet test the username, password, SSO token, group, or subsequent firewall rule.

Check the authentication method per service

Under Authentication > Services, each service has its own server selection. Therefore, don’t just check whether an AD, LDAP, RADIUS, or Entra server exists. Verify that it is selected for the affected service:

  • Firewall authentication methods for user traffic and classic firewall authentication;
  • User portal authentication methods for the User Portal;
  • VPN portal authentication methods for the VPN Portal;
  • VPN (IPsec/dial-in/L2TP/PPTP) authentication methods for these VPN sign-ins;
  • Administrator authentication methods for named administrators;
  • SSL VPN authentication methods for SSL VPN.

With multiple servers, the order also matters. Don’t move a server to the top prematurely. First document the current order, default group, and working comparison user. A change could otherwise affect other users or services. For external administrators, TACACS+ for Sophos Firewall WebAdmin also separates server authentication, the local administrator profile, and a lockout-safe fallback.

Interpret Test connection correctly

Under Authentication > Servers, Test connection checks the credentials and the connection between the firewall and server. If it fails, first correct DNS, routing, port, connection security, certificate chain, service account, or password.

A successful test, however, is only the first layer. For AD SSO, it doesn’t test Kerberos, NTLM, the SPN, redirection location, browser trust, or actual user access. For portal and VPN sign-ins, the service selection, group, MFA, and policy must also be tested separately. Connect Active Directory to Sophos Firewall describes the complete AD baseline.

Use Live Users as the first identity proof

After a fresh sign-in, open Current activities > Live users. At least three values matter:

For the Client Authentication Agent, the client type must be Authentication agent. Also check the path to 1.2.3.4:9922, the Authentication Server CA, and the actual firewall rule match separately.

  • User: Is the detected user really the pilot account?
  • IP address: Does the address match the observed client connection?
  • Client Type: Was the expected path used, such as AD SSO Kerberos, AD SSO NTLM, STAS, Multi-host client, Captive portal, RADIUS SSO, or a VPN type?

If the user is missing, the identity layer hasn’t succeeded. Don’t continue troubleshooting the later user rule. If a computer name appears instead of the signed-in user, an operating-system service such as Windows Update may have sent the device’s NTLM credentials. An actual interactive user sign-in and a new web request isolate this case.

An external AD, LDAP, or RADIUS user normally appears under Authentication > Users only after their first successful sign-in to a firewall service. A missing local user record therefore doesn’t prove that the directory account is missing. For an account maintained directly on SFOS, create and manage normal local users instead provides the complete account, service, testing, and offboarding workflow.

For a clean repeat test, a normal existing session can be disconnected under Live users. After an AD SSO session is manually disconnected, it can take up to three minutes before the user can sign in again. Clientless users aren’t signed out there with Disconnect; set them to Inactive under Authentication > Clientless users instead.

Check the user object, main group, and restrictions

When the user is visible or a local record already exists, continue under Authentication > Users. Don’t change values blindly. First document the effective properties:

  • Active or Inactive status;
  • Group field as the main group for AD users;
  • Other group memberships;
  • user-specific policy overrides;
  • access time and sign-in restriction;
  • surfing quota and network traffic quota;
  • MFA assignment for the affected service;
  • additional User ID property.

Don’t treat the main group and other groups as equivalent

An AD user can belong to several groups. Sophos Firewall shows the effective main group in the Group field and other memberships separately. Firewall rules, web policies, and some other features can consider multiple groups. Access time, quotas, MFA, and several remote-access features, however, use only the main group or an explicit user assignment.

Therefore, don’t change the group order as a quick fix. First establish which function fails and which group logic it supports. After a planned change, authenticate the user again and recheck the main group.

The complete operational workflow from creating groups and choosing the Default Group through Reorder, multiple groups, user overrides, and rollback is described in Manage Sophos Firewall user groups and the main group correctly.

Treat access time, quotas, and MFA as separate causes

An unexpected Captive Portal or rejected sign-in may also be caused by a blocked access time, an exhausted surfing or network traffic quota, or missing MFA enrollment. Compare these settings on the user and their effective group.

The respective checks and recovery paths are described in Access Time for users and groups, Surfing and Network Traffic Quota, and MFA for Sophos Firewall. Don’t reset a quota or broaden an access-time policy before proving its effective assignment.

Attribute the User ID limit only after seeing the value

Under Show additional properties, the User ID can be displayed. A value above 65535 is outside the supported range; the user then can’t authenticate or become a live user. The number of visible users alone isn’t enough for this diagnosis because groups share the same ID range.

If the ID is valid or the user is still completely absent, continue the general diagnosis. Purge AD users belongs only in a proven cleanup case. The safe specialist workflow is in Check the Sophos Firewall User ID limit.

Continue with the specific method in use

Once the service and failure layer are known, investigate only the relevant method path. This keeps troubleshooting small and avoids changing several authentication systems at once.

Classic sign-in and Captive Portal

For a manual sign-in, check the direct portal URL, Device Access, selected authentication method, user status, password, and MFA. For Captive Portal, also check whether the unknown connection reaches the intended rule with Use web authentication for unknown users and whether DNS works before sign-in.

The full setup, including port 8090, Live Users, the user rule, and sign-out, is described in Set up and test Sophos Firewall Captive Portal.

AD SSO with Kerberos or NTLM

A successful AD connection test isn’t enough for SSO. The hostname or FQDN, DNS resolution, redirection location, certificate, browser trust, domain join, and, for Kerberos, the matching HTTP SPN must also be correct. If Kerberos falls back to NTLM or Captive Portal, check this naming and trust path first.

nasm.log is relevant to NTLM and Kerberos errors. Don’t blindly change the SPN or domain join; first capture the existing state on the firewall and domain controller.

STAS, SATC, and multi-user hosts

With Synchronized User ID Authentication, the Windows domain identity and client IP arrive through Security Heartbeat. The endpoint status, UPN domain, sAMAccountName, AD validation, and Live users must then match; a green heartbeat alone doesn’t prove the user rule.

With STAS, the same client IP must be tracked through Windows events, STA Agent, Collector, and firewall. If the user is only missing from Live Users, check Collector reachability, monitored networks, the exclusion list, and Client Authentication. The full workflow is in Set up STAS on Sophos Firewall.

Several concurrent users behind one terminal-server IP require a session model. SATC for Remote Desktop Services can associate multiple protocols with a session. Per-Connection AD SSO for multi-user hosts, by contrast, applies only to HTTP and HTTPS connections through the Direct Web Proxy. Normal IP-based mapping can’t reliably distinguish several users behind the same address.

RADIUS SSO, Entra ID, and VPN

RADIUS SSO requires accounting events with the correct client IP; a successful RADIUS access test doesn’t prove this accounting path. RADIUS SSO with accounting describes how to configure and verify Framed-IP-Address.

For Microsoft Entra ID SSO, also check the service-specific OAuth flow. Captive Portal, VPN Portal, and WebAdmin use different redirect URIs, authentication methods, and log files. A successful sign-in to one of these services doesn’t prove the others.

For VPN, after authentication also verify that the user or their group is in the correct remote-access policy. A green tunnel status doesn’t yet prove access to internal destinations.

Authentication succeeds but traffic is still blocked

When the user is visible under Live users with the expected IP, move the diagnosis to authorization and the packet path. A real test flow must show:

  1. Which Firewall Rule ID processes the connection?
  2. Does this rule contain the expected user or a supported group?
  3. Is Log firewall traffic turned on?
  4. Does the rule select the correct web, application, IPS, or traffic-shaping policy?
  5. Do the destination, service, NAT, route, and return path match?

A general IP rule above the user rule may already process the traffic. Likewise, the correct user rule may match while web policy, TLS inspection, DNS, routing, or the destination system blocks later access. Therefore, don’t use an authentication change as the fix for a proven routing or policy error.

Test a Sophos Firewall rule correctly explains the complete validation using Log Viewer, Policy tester, and Packet Capture.

Correlate Log Viewer and log files

In Log viewer, check the Authentication module using the recorded user, source IP, and narrow test period. Then find the same time in the firewall or VPN module. This shows whether sign-in itself fails or only the subsequent data flow is blocked.

For classic authentication, authorization, and accounting, access_server.log is the first file check. In the Advanced Shell, the following commands are read-only:

cd /log
tail -n 200 access_server.log

Add only the file relevant to the method:

tail -n 200 nasm.log
tail -n 200 oauth_sso_captive.log
tail -n 200 oauth_sso_webadmin.log
tail -n 200 oauth_sso_vpn.log
  • nasm.log is for NTLM and Kerberos.
  • oauth_sso_captive.log is for the Entra SSO flow of Captive Portal.
  • oauth_sso_webadmin.log is for Entra SSO to WebAdmin.
  • oauth_sso_vpn.log is for Entra SSO to VPN Portal, IPsec, and SSL VPN.

The commands only read the last 200 lines. They don’t enable debug or restart a service. Add files such as vpnportal.log, sslvpn.log, strongswan.log, chromebook-sso-backend.log, or csd.log only after selecting the method. Sophos Firewall service logs explains the mapping.

In an HA cluster, each node stores logs and reports only for traffic it processes. Record the role, test time, and processing node, and check the files on that node. An empty file on the other device doesn’t disprove the fault.

Troubleshoot by symptom

Test connection succeeds, but SSO doesn’t work

The authentication server and its credentials are reachable. Next, check Authentication > Services, Device Access, Live Users, and the method-specific SSO path. For AD, this especially includes DNS, FQDN, redirection location, SPN, browser trust, domain join, and nasm.log.

Captive Portal appears instead of transparent sign-in

First determine whether AD SSO or STAS should identify the user before web access. Then check Client Type under Live Users, the Authentication module, access time, quotas, and credentials. The portal is a symptom of a missing or rejected identity, not automatically a portal fault.

Computer name appears instead of username

A system service may send the endpoint’s NTLM credentials even though no user is interactively signed in. Sign in an actual user, open a new browser session, and repeat the test. If the computer name remains, check the browser or proxy path, SSO method, and nasm.log.

User is missing from Live Users

First check the service selection, Device Access, and the specific method. For STAS, the Collector and firewall must see the same IP; for Captive Portal, sign-in must actually complete; for external servers, the local user record is only created after successful sign-in. Check the User ID, sign-in restriction, and quotas only when a corresponding user record exists.

User is visible, but the group or policy is wrong

Under Authentication > Users, compare the main group, other group memberships, and user overrides. Then verify whether the specific feature supports multiple groups. Change group order only in a planned maintenance window because it may also shift MFA, quotas, access time, VPN, and other policies.

Sign-in succeeds, but internet or the VPN destination remains unreachable

Authentication is no longer the first fault layer. Check the Firewall Rule ID, rule order, logging, web or VPN policy, NAT, route, and return path using a real flow. Both user identity and packet path must be correct.

The fault occurs only after HA failover or intermittently

Record the firmware build, role change, active node role, sign-in time, and authentication method. Test a new sign-in and connection after failover. Then check the logs on the node that processed the attempt. Don’t treat an existing session as proof of a new successful authentication.

Stop conditions and support data

Don’t make further changes to the production configuration when:

  • the affected service or authentication method isn’t unambiguous;
  • Test connection fails while DNS, route, port, TLS, or credentials remain unresolved;
  • the pilot user appears with an unexpected client type or wrong source IP;
  • the main group, user override, or quota can’t be explained;
  • a negative test unexpectedly obtains access;
  • only broad Device Access or a broad firewall rule appears to make the test work;
  • an HA problem can’t be assigned to the processing node and time.

Before escalation, capture at least:

  • appliance model, SFOS version, and full build;
  • for HA, both roles and the processing node;
  • authentication source and affected service;
  • test user, source IP, client type, and time window;
  • server order under Authentication > Services;
  • Test connection result with its limited meaning;
  • status, main group, other groups, overrides, and User ID;
  • authentication event, Firewall Rule ID, and outcome of the real test flow;
  • matching authentication, portal, VPN, and firewall logs.

Usernames and groups may contain sensitive information. Provide the package only through the intended support channel. Create a Consolidated Troubleshooting Report and targeted log export before service restarts, debug, or data cleanup.

Checklist

  • The test case contains the user, source IP, service, time, and expected result.
  • The local service is reachable only from the intended zone or source.
  • The correct server is selected under Authentication > Services for the affected service.
  • Test connection wasn’t mistaken for a complete SSO test.
  • Live Users shows the user, IP, and expected client type.
  • Status, main group, other groups, overrides, and User ID have been checked.
  • Access time, quotas, sign-in restriction, and MFA were evaluated only on the effective user path.
  • Authentication success and the subsequent firewall, web, or VPN flow were tested separately.
  • Log Viewer and the matching log file were correlated over the same time window.
  • In HA, the processing node was checked.
  • No broad restart, purge, or broad allow was used as a quick fix.
  • Support data was captured before further changes.

Frequently asked questions

Why doesn't SSO work even though Test connection succeeds?

Test connection only checks credentials and reachability of the authentication server. SSO also needs the correct service assignment, Device Access, and, depending on the method, DNS, SPN, browser trust, redirect URI, accounting, or an agent or Collector path. Only an actual user sign-in with the matching client type under Live Users tests this workflow.

Why does a live user still have no access?

Live Users confirms the detected identity and source IP, but not authorization or the data path. Check the main group, user override, access time, quota, Firewall Rule ID, web or VPN policy, NAT, routing, and return path separately.

Which log file should I check first for an authentication error?

For classic authentication, authorization, and accounting, start with access_server.log. Add nasm.log for NTLM or Kerberos. Depending on the service, Entra SSO uses oauth_sso_captive.log, oauth_sso_webadmin.log, or oauth_sso_vpn.log. The previously recorded test time is always essential.