Systematically troubleshoot Sophos Firewall authentication errors
An authentication error on Sophos Firewall can occur at very different points. The server may be reachable but not selected for the affected service. Sign-in may work while the main group, quota, or firewall rule prevents access. Or the firewall may not see the user at all and processes the traffic only by IP address.
The safe quick path deliberately separates these layers:
- Record the user, source IP, affected service, time, and expected result.
- Test reachability of the portal or authentication service from the actual source zone.
- Under Authentication > Services, verify that the correct server is selected for the affected service and is in the intended position.
- Run Test connection on the authentication server, but treat it only as proof of credentials and server connectivity.
- Generate a fresh sign-in and check the user, source IP, and Client Type under Current activities > Live users.
- Under Authentication > Users, check status, main group, other groups, user overrides, and User ID.
- Check access time, quotas, MFA, and sign-in restrictions only for the user actually affected or their effective group.
- Only after identity succeeds, test the expected firewall rule, Rule ID, web or VPN policy, NAT, and routing with real traffic.
⚠️ Don’t broadly change authentication servers, group order, quotas, or Device Access on suspicion. A service restart, Purge AD users, or an Any rule isn’t a normal first diagnostic step. First prove at which layer the workflow stops.
Five layers instead of one sign-in error
In practice, a simple model helps. User access doesn’t pass just one check but at least five separate layers:
- Reachability: The client and firewall reach the portal, proxy, or authentication server through the intended path.
- Service and method: The correct authentication server is selected under Authentication > Services for this exact service.
- Identity: The firewall recognizes the expected user and shows them with the correct source IP and client type.
- Authorization: User status, group, main group, access time, quota, MFA, and service-specific permissions match.
- Subsequent traffic: A firewall rule, web policy, or VPN policy allows the required destinations, and the return path works.
This separation prevents two common false conclusions. A green Test connection doesn’t prove that SSO sign-in works. Conversely, a user under Live users doesn’t yet prove that their traffic matches the intended rule or policy.
Define a reproducible test case
Before the first change, define one test case. Safe documentation values include:
- User:
auth-pilot@example.com - Client IP:
192.0.2.25 - Source zone:
LAN - Service:
Firewall authentication,User portal,VPN portal,SSL VPN, orWeb admin console - Test time: local firewall time with date and minute
- Expected group:
Internet-Standard - Expected rule:
LAN-Users-to-WAN - Expected result: successful sign-in and HTTPS access to a defined test destination
example.com and 192.0.2.0/24 are documentation values. Replace the username, source IP, zone, group, rule, and destination with real pilot values. The pilot should use the same authentication source and permission group as the affected user but have no unnecessary administrative privileges.
If possible, also test a working comparison user from the same source. If both fail, the cause is more likely the service, server, or reachability. If only one user fails, the user object, groups, quotas, MFA, or sign-in restrictions are more likely.
Check the service, reachability, and server
Check Device Access only for the required service
Portals and local authentication services aren’t allowed by normal firewall rules. Under Administration > Device access, or with a targeted Local service ACL exception rule, the required service must be reachable from the intended source zone or known source networks.
Different entries may be relevant depending on the test, such as Captive portal, User portal, VPN portal, SSL VPN, AD SSO, or Client Authentication. Check only the service actually required. Broadly allowing LAN or WAN would hide the fault and may expose additional management or portal services. Device Access and Local Service ACL on Sophos Firewall explains the safe boundary.
Reachability alone isn’t authentication. A visible portal page only confirms that the client reaches the local service. It doesn’t yet test the username, password, SSO token, group, or subsequent firewall rule.
Check the authentication method per service
Under Authentication > Services, each service has its own server selection. Therefore, don’t just check whether an AD, LDAP, RADIUS, or Entra server exists. Verify that it is selected for the affected service:
- Firewall authentication methods for user traffic and classic firewall authentication;
- User portal authentication methods for the User Portal;
- VPN portal authentication methods for the VPN Portal;
- VPN (IPsec/dial-in/L2TP/PPTP) authentication methods for these VPN sign-ins;
- Administrator authentication methods for named administrators;
- SSL VPN authentication methods for SSL VPN.
The firewall supports no more than 20 authentication servers in total, and you can select no more than 20 servers for each authentication method. If another server is missing when adding it or from the selection, this isn’t automatically a connection failure. First check the total count, the selection, and whether every server is still required.
The User Portal and VPN Portal can inherit the firewall authentication server list through Set authentication methods same as firewall. SSL VPN can likewise use Same as VPN or Same as firewall. This inheritance is convenient, but it extends the effect of later changes to server selection or order. Always document the effective inherited list during diagnosis. Administrator authentication methods also do not apply to the local super administrator admin; retain and test that recovery path separately.
Check global session limits separately: Maximum session timeout ends a successfully authenticated session after its maximum duration. SFOS checks authorization every three minutes; access policies, surfing quota, data transfer limit, and the maximum session length can restrict access earlier. The global Simultaneous logins value only applies to users created after the value was set.
In SFOS 22 and SFOS 23, also check the method-specific values under Authentication > Services separately: AD SSO settings (NTLM and Kerberos) and Web client settings (iOS, Android and API) each have Inactivity time and Data transfer threshold settings. If less than the minimum amount of data is transferred within the configured period, the user is considered inactive; Inactivity time determines how long the user must be inactive before being signed out. If a user is unexpectedly signed out, first document the Client Type used, both values in the relevant settings block, and the actual traffic over the same period. Don’t change the global Maximum session timeout or STAS/Captive Portal timers instead.
With multiple servers, the order also matters. Don’t move a server to the top prematurely. First document the current order, default group, and working comparison user. A change could otherwise affect other users or services. For external administrators, TACACS+ for Sophos Firewall WebAdmin also separates server authentication, the local administrator profile, and a lockout-safe fallback.
Replace eDirectory before upgrading: The native eDirectory path belongs to SFOS 22. Before upgrading to SFOS 23.0 or later, validate a supported replacement for every affected service and remove the native eDirectory configuration; retaining it causes the upgrade to fail. Don’t delete production authentication before verifying target sign-ins, groups, policies, and an independent local fallback. Migrate eDirectory safely describes the controlled cutover and recovery path.
Interpret Test connection correctly
Under Authentication > Servers, Test connection checks the credentials and the connection between the firewall and server. If it fails, first correct DNS, routing, port, connection security, certificate chain, service account, or password.
A successful test, however, is only the first layer. For AD SSO, it doesn’t test Kerberos, NTLM, the SPN, redirection location, browser trust, or actual user access. For portal and VPN sign-ins, the service selection, group, MFA, and policy must also be tested separately. Connect Active Directory to Sophos Firewall describes the complete AD baseline.
Use Live Users as the first identity proof
After a fresh sign-in, open Current activities > Live users. At least three values matter:
Capacity limit: SFOS 22 supports up to
32,000live users. Users and groups share the User ID range; for diagnosis, a visible ID through65535is valid. In very large environments, check this limit separately from an individual failed sign-in. Don’t delete sessions or user objects indiscriminately.
For the Client Authentication Agent, the client type must be Authentication agent. Also check the path to 1.2.3.4:9922, the Authentication Server CA, and the actual firewall rule match separately.
- User: Is the detected user really the pilot account?
- IP address: Does the address match the observed client connection?
- Client Type: Was the expected path used, such as
AD SSO Kerberos,AD SSO NTLM,STAS,Multi-host client,Captive portal,RADIUS SSO, or a VPN type?
Interpret the type by its source rather than as one generic sign-in. Browser-based captive portal and guest sessions appear as Web client, Android web client, or iOS web client. Client-driven mappings include STAS, Thin client, Heartbeat, and API; server-driven SSO includes RADIUS SSO and Chromebook SSO. eDirectory SSO belongs only to the SFOS 22 list; SFOS 23 removes native eDirectory SSO and no longer lists this type. VPN entries distinguish IPsec VPN, SSL VPN, L2TP VPN, and PPTP VPN. The observed type must match the intended mechanism.
Clientless users are the exception to the fresh sign-in test: the list shows them as soon as they’re configured because SFOS authenticates them by IP address, and it omits them when they’re inactive.
If the user is missing, the identity layer hasn’t succeeded. Don’t continue troubleshooting the later user rule. If a computer name appears instead of the signed-in user, an operating-system service such as Windows Update may have sent the device’s NTLM credentials. An actual interactive user sign-in and a new web request isolate this case.
An external AD, LDAP, or RADIUS user normally appears under Authentication > Users only after their first successful sign-in to a firewall service. A missing local user record therefore doesn’t prove that the directory account is missing. For an account maintained directly on SFOS, create and manage normal local users instead provides the complete account, service, testing, and offboarding workflow.
For a clean repeat test, select a normal existing session under Live users, click Disconnect, optionally change the notification text, and confirm Disconnect. After an AD SSO session is manually disconnected, it can take up to three minutes before the user can sign in again. The notification reaches only users signed in with Authentication agent, Android client, iOS client, or Chromebook SSO.
Clientless users aren’t signed out with Disconnect. Set the specific user to Inactive under Authentication > Clientless users instead. If Disconnect has already been used on a clientless user and access must be restored, change the user’s status to Inactive and then to Active.
Check the user object, main group, and restrictions
When the user is visible or a local record already exists, continue under Authentication > Users. Don’t change values blindly. First document the effective properties:
- Active or Inactive status;
- Group field as the main group for AD users;
- Other group memberships;
- user-specific policy overrides;
- access time and sign-in restriction;
- surfing quota and network traffic quota;
- MFA assignment for the affected service;
- additional User ID property.
Don’t treat the main group and other groups as equivalent
An AD user can belong to several groups. Sophos Firewall shows the effective main group in the Group field and other memberships separately. Firewall rules, web policies, and some other features can consider multiple groups. Access time, quotas, MFA, and several remote-access features, however, use only the main group or an explicit user assignment.
Therefore, don’t change the group order as a quick fix. First establish which function fails and which group logic it supports. After a planned change, authenticate the user again and recheck the main group.
The complete operational workflow from creating groups and choosing the Default Group through Reorder, multiple groups, user overrides, and rollback is described in Manage Sophos Firewall user groups and the main group correctly.
Treat access time, quotas, and MFA as separate causes
An unexpected Captive Portal or rejected sign-in may also be caused by a blocked access time, an exhausted surfing or network traffic quota, or missing MFA enrollment. Compare these settings on the user and their effective group.
The respective checks and recovery paths are described in Access Time for users and groups, Surfing and Network Traffic Quota, and MFA for Sophos Firewall. Don’t reset a quota or broaden an access-time policy before proving its effective assignment.
If only MFA fails, check Issued tokens, the token status, and the firewall and phone clocks. Synchronize token time offset checks and corrects clock drift. SFOS 22 supports SHA1, SHA256, and SHA512; the authenticator app must support the configured algorithm. Changing the algorithm or deleting a token isn’t a diagnostic test but a planned migration or re-enrollment.
Attribute the User ID limit only after seeing the value
Under Show additional properties, the User ID can be displayed. A value above 65535 is outside the supported range; the user then can’t authenticate or become a live user. The number of visible users alone isn’t enough for this diagnosis because groups share the same ID range.
If the ID is valid or the user is still completely absent, continue the general diagnosis. Purge AD users belongs only in a proven cleanup case. The safe specialist workflow is in Check the Sophos Firewall User ID limit.
Continue with the specific method in use
Once the service and failure layer are known, investigate only the relevant method path. This keeps troubleshooting small and avoids changing several authentication systems at once.
Classic sign-in and Captive Portal
For a manual sign-in, check the direct portal URL, Device Access, selected authentication method, user status, password, and MFA. For Captive Portal, also check whether the unknown connection reaches the intended rule with Use web authentication for unknown users and whether DNS works before sign-in. The username and password can each contain no more than 50 characters.
The full setup, including port 8090, Live Users, the user rule, and sign-out, is described in Set up and test Sophos Firewall Captive Portal.
AD SSO with Kerberos or NTLM
A successful AD connection test isn’t enough for SSO. The hostname or FQDN, DNS resolution, redirection location, certificate, browser trust, domain join, and, for Kerberos, the matching HTTP SPN must also be correct. If Kerberos falls back to NTLM or Captive Portal, check this naming and trust path first.
nasm.log is relevant to NTLM and Kerberos errors. Don’t blindly change the SPN or domain join; first capture the existing state on the firewall and domain controller.
If /log/nasm.log shows a KVNO error, the client may still hold a Kerberos ticket with the old key version after the firewall has rejoined the domain. First lock and unlock the Windows client or sign the user out and back in so that it obtains a new ticket. If the error remains, check the Key Version Number, firewall computer object, and SPN on the domain controller; another domain join isn’t the first step.
STAS, SATC, and multi-user hosts
With Synchronized User ID Authentication, the Windows domain identity and client IP arrive through Security Heartbeat. The endpoint status, UPN domain, sAMAccountName, AD validation, and Live users must then match; a green heartbeat alone doesn’t prove the user rule.
With STAS, the same client IP must be tracked through Windows events, STA Agent, Collector, and firewall. If the user is only missing from Live Users, check Collector reachability, monitored networks, the exclusion list, and Client Authentication. The full workflow is in Set up STAS on Sophos Firewall.
The Legacy SATC Client is no longer supported. Sophos recommends migrating SATC installations to Sophos Server Protection in Sophos Fusion. Before troubleshooting, establish which SATC client is actually installed; changing firewall rules or groups won’t repair an obsolete legacy client.
Several concurrent users behind one terminal-server IP require a session model. SATC for Remote Desktop Services can associate multiple protocols with a session. Per-Connection AD SSO for multi-user hosts, by contrast, applies only to HTTP and HTTPS connections through the Direct Web Proxy. Normal IP-based mapping can’t reliably distinguish several users behind the same address.
RADIUS SSO, Entra ID, and VPN
RADIUS SSO requires accounting events with the correct client IP; a successful RADIUS access test doesn’t prove this accounting path. RADIUS SSO with accounting describes how to configure and verify Framed-IP-Address. The VPN Portal doesn’t support RADIUS authentication using challenge-based MFA, so a rejected challenge flow there isn’t a server-order fault.
SFOS 22: For Microsoft Entra ID SSO, also check the service-specific OAuth flow. Captive Portal, VPN Portal, and WebAdmin use different redirect URIs, authentication methods, and log files. A successful sign-in to one of these services doesn’t prove the others.
SFOS 23.0: Under Authentication > Servers, the server type is OpenID Connect; IdP vendor distinguishes Microsoft Entra ID from Google Workspace. Use the versioned validation paths for the actual service: Entra WebAdmin, Entra VPN Portal and remote access, or Google Workspace OIDC. Redirect URI, service assignment, groups or administrator mapping, MFA, and positive, negative, and local fallback tests remain separate checks. A successful IdP sign-in doesn’t automatically grant administrator rights.
For VPN, after authentication also verify that the user or their group is in the correct remote-access policy. A green tunnel status doesn’t yet prove access to internal destinations.
Authentication succeeds but traffic is still blocked
When the user is visible under Live users with the expected IP, move the diagnosis to authorization and the packet path. A real test flow must show:
- Which Firewall Rule ID processes the connection?
- Does this rule contain the expected user or a supported group?
- Is Log firewall traffic turned on?
- Does the rule select the correct web, application, IPS, or traffic-shaping policy?
- Do the destination, service, NAT, route, and return path match?
A general IP rule above the user rule may already process the traffic. Likewise, the correct user rule may match while web policy, TLS inspection, DNS, routing, or the destination system blocks later access. Therefore, don’t use an authentication change as the fix for a proven routing or policy error.
Test a Sophos Firewall rule correctly explains the complete validation using Log Viewer, Policy tester, and Packet Capture.
Correlate Log Viewer and log files
In Log viewer, check the Authentication module using the recorded user, source IP, and narrow test period. Then find the same time in the firewall or VPN module. This shows whether sign-in itself fails or only the subsequent data flow is blocked.
For classic authentication, authorization, and accounting, access_server.log is the first file check. In the Advanced Shell, the following commands are read-only:
cd /log
tail -n 200 access_server.log
Add only the file relevant to the method:
tail -n 200 nasm.log
The following service-specific Entra read-only examples apply only to SFOS 22, not SFOS 23:
tail -n 200 oauth_sso_captive.log
tail -n 200 oauth_sso_webadmin.log
tail -n 200 oauth_sso_vpn.log
nasm.logis for NTLM and Kerberos.- SFOS 22:
oauth_sso_captive.logis for the Entra SSO flow of Captive Portal. - SFOS 22:
oauth_sso_webadmin.logis for Entra SSO to WebAdmin. - SFOS 22:
oauth_sso_vpn.logis for Entra SSO to VPN Portal, IPsec, and SSL VPN.
SFOS 23.0: Check the OIDC sign-in flow in /log/oauth_sso_svc.log. For Entra, a TLS error from Test connection instead belongs in /log/sfos-macro-cfg.log. In Log viewer, use Admin for WebAdmin and Authentication for Captive Portal, VPN Portal, IPsec, and SSL VPN. The provider workflows linked above match validation and troubleshooting to the actual service; no restart is required for this read-only diagnosis.
The commands only read the last 200 lines. They don’t enable debug or restart a service. Add files such as vpnportal.log, sslvpn.log, strongswan.log, chromebook-sso-backend.log, or csd.log only after selecting the method. Sophos Firewall service logs explains the mapping.
In an HA cluster, each node stores logs and reports only for traffic it processes. Record the role, test time, and processing node, and check the files on that node. An empty file on the other device doesn’t disprove the fault.
Troubleshoot by symptom
Test connection succeeds, but SSO doesn’t work
The authentication server and its credentials are reachable. Next, check Authentication > Services, Device Access, Live Users, and the method-specific SSO path. For AD, this especially includes DNS, FQDN, redirection location, SPN, browser trust, domain join, and nasm.log.
Captive Portal appears instead of transparent sign-in
First determine whether AD SSO or STAS should identify the user before web access. Then check Client Type under Live Users, the Authentication module, access time, quotas, and credentials. The portal is a symptom of a missing or rejected identity, not automatically a portal fault.
Computer name appears instead of username
A system service may send the endpoint’s NTLM credentials even though no user is interactively signed in. Sign in an actual user, open a new browser session, and repeat the test. If the computer name remains, check the browser or proxy path, SSO method, and nasm.log.
User is missing from Live Users
First check the service selection, Device Access, and the specific method. For STAS, the Collector and firewall must see the same IP; for Captive Portal, sign-in must actually complete; for external servers, the local user record is only created after successful sign-in. Check the User ID, sign-in restriction, and quotas only when a corresponding user record exists.
User is visible, but the group or policy is wrong
Under Authentication > Users, compare the main group, other group memberships, and user overrides. Then verify whether the specific feature supports multiple groups. Change group order only in a planned maintenance window because it may also shift MFA, quotas, access time, VPN, and other policies.
Sign-in succeeds, but internet or the VPN destination remains unreachable
Authentication is no longer the first fault layer. Check the Firewall Rule ID, rule order, logging, web or VPN policy, NAT, route, and return path using a real flow. Both user identity and packet path must be correct.
The fault occurs only after HA failover or intermittently
Record the firmware build, role change, active node role, sign-in time, and authentication method. Test a new sign-in and connection after failover. Then check the logs on the node that processed the attempt. Don’t treat an existing session as proof of a new successful authentication.
Stop conditions and support data
Don’t make further changes to the production configuration when:
- the affected service or authentication method isn’t unambiguous;
- Test connection fails while DNS, route, port, TLS, or credentials remain unresolved;
- the pilot user appears with an unexpected client type or wrong source IP;
- the main group, user override, or quota can’t be explained;
- a negative test unexpectedly obtains access;
- only broad Device Access or a broad firewall rule appears to make the test work;
- an HA problem can’t be assigned to the processing node and time.
Before escalation, capture at least:
- appliance model, SFOS version, and full build;
- for HA, both roles and the processing node;
- authentication source and affected service;
- test user, source IP, client type, and time window;
- server order under Authentication > Services;
- Test connection result with its limited meaning;
- status, main group, other groups, overrides, and User ID;
- authentication event, Firewall Rule ID, and outcome of the real test flow;
- matching authentication, portal, VPN, and firewall logs.
Usernames and groups may contain sensitive information. Provide the package only through the intended support channel. Create a Consolidated Troubleshooting Report and targeted log export before service restarts, debug, or data cleanup.
Checklist
- The test case contains the user, source IP, service, time, and expected result.
- The local service is reachable only from the intended zone or source.
- The correct server is selected under Authentication > Services for the affected service.
- Test connection wasn’t mistaken for a complete SSO test.
- Live Users shows the user, IP, and expected client type.
- Status, main group, other groups, overrides, and User ID have been checked.
- Access time, quotas, sign-in restriction, and MFA were evaluated only on the effective user path.
- Authentication success and the subsequent firewall, web, or VPN flow were tested separately.
- Log Viewer and the matching log file were correlated over the same time window.
- In HA, the processing node was checked.
- No broad restart, purge, or broad allow was used as a quick fix.
- Support data was captured before further changes.
Frequently asked questions
Why doesn't SSO work even though Test connection succeeds?
Why does a live user still have no access?
Which log file should I check first for an authentication error?
access_server.log. Add nasm.log for NTLM or Kerberos. SFOS 22 only: Depending on the service, Entra SSO uses oauth_sso_captive.log, oauth_sso_webadmin.log, or oauth_sso_vpn.log. For SFOS 23, use the OIDC log branch and the provider and module selection described above. The previously recorded test time is always essential.