In 20 years, you will be more dissapointed by what you didn't do than by what you did.

Palo Alto PBF Failover Not Working: Dual-ISP Troubleshooting Checklist

A backup internet circuit is not useful if the firewall keeps selecting a broken path—or switches paths but the application still cannot reconnect. Use this Palo Alto policy-based forwarding (PBF) checklist to separate rule selection, path monitoring, fallback routing, NAT and return-path faults before changing production policy.

This is an original diagnostic worksheet for a hypothetical dual-ISP branch, not a recorded lab result. Product behavior is anchored to the PAN-OS 11.1 documentation and a Palo Alto LIVEcommunity technical article; validate it against your installed release. It is not a Prisma SD-WAN or firewall HA configuration guide.

Why the routing table can look correct while traffic goes the wrong way

PBF can select an alternative path to the next hop in the routing table, including a different egress interface.[2] Therefore, start with the actual affected flow, not just a screenshot of the default route. A correct fallback route does not by itself prove that the production session is using it.

Record the test client, destination, protocol, destination port, start time, session identifier and whether the connection existed before the failure. Keep these identifiers in a private incident worksheet rather than a public troubleshooting post. Use one explicitly approved test destination and one repeatable application transaction.

The most useful first question is: Does a newly created connection fail, an already established connection fail, or both? Palo Alto's path-monitoring documentation explicitly distinguishes established sessions from new sessions and also distinguishes a rule that stays enabled from one disabled by monitoring.[1] Do not turn a successful new browser connection into a claim that every existing application survived.

Dual-ISP flow with separate preferred and backup paths through PBF, NAT and application validation
Source: original Network freak diagram; hypothetical generic topology, not a customer network.

Troubleshooting matrix: choose the next test, not the next guess

The following is a recommended investigation sequence. Its observations are clues, not conclusive diagnoses.

Observed symptom Compare first Evidence to collect Narrow next action
Traffic uses ISP A although the ordinary route points to ISP B Matching PBF rule versus routing decision Flow tuple, matching policy and observed egress Inspect rule order and match criteria before editing routes
ISP A is unusable but its monitor remains up Probe target versus failing service Monitor target, reachability and application transaction Determine whether the target can still answer during the actual outage
Monitor is down but the backup connection fails Fallback policy/routing versus ISP B translation and return path Egress, translated source, reply packets and session state Validate the complete backup path, not only monitor status
New sessions work, existing application stalls Fresh connection versus pre-failure session Separate timestamps and session identifiers Investigate reconnect behavior and release-specific monitoring action
Only inbound access through ISP B fails Request ingress versus reply egress Captures in both directions and return-MAC information where applicable Evaluate symmetric return on the relevant inbound flow
Traffic alternates repeatedly between circuits Monitor transitions versus application failures Timestamped probe, link and application events Investigate target reliability before making detection more aggressive
Firewall-originated ping works but client traffic fails Device-originated probe versus client data flow Source context, interfaces, policy and NAT for each Repeat the test from the actual client path

Do not change PBF, NAT and security policy simultaneously. That may restore service without revealing which change mattered, leaving a fragile rollback plan. Start with read-only observation, propose one bounded change, and compare the same transaction afterward.

Path monitoring: decide exactly what failure you are detecting

PBF path monitoring uses ICMP heartbeats to test reachability to an IP address, with a monitoring profile defining the failure threshold.[1] An ICMP answer is not an application-readiness test. Treat it as evidence about the chosen target, not proof that DNS, TLS or the business transaction works.

For the hypothetical branch, choose a target whose location matches the intended failure domain. A directly connected gateway may be useful for testing local next-hop reachability; a controlled target beyond it may reveal a wider upstream failure. Neither choice is universally correct. Write down what an answering target proves, what it does not prove, who owns it, and whether ICMP filtering or rate limiting could invalidate your design.

Review the monitoring profile action and the rule's disable-on-monitor-failure setting together. The vendor documents fail-over and wait-recover actions, continued monitoring, and return to the original route after recovery.[1] The exact established-session behavior belongs in your release-specific test plan, not in an assumption inferred from the word “fail-over.” Consult the original behavior table for your PAN-OS release; flattened copies of a multi-column table are easy to misread.

If you disable a rule on monitor failure, new sessions can be evaluated against remaining PBF rules before ordinary routing is used.[1] Inspect those remaining rules. A lower-priority catch-all rule can be just as important to your investigation as the intended backup default route.

Do not promise session survival from a path change

In a conventional dual-ISP design, changing the translated public source address changes the identity of the connection as seen by the remote endpoint. Plan for application reconnection rather than promising transparent continuity. Independently test the application recovery objective: a working routing decision is not the same thing as a successful login, upload or long-running job.

This is also why a single ping is a weak acceptance test. Keep a pre-existing test connection running and create fresh connections during the failure. Label their results separately. Do not clear all production sessions to make a failover demonstration look clean.

Implementation worksheet: define the backup path before editing policy

Use this table as a change-review asset. Fill it in privately for both circuits, and get agreement on the rollback owner and recovery target before the maintenance window.

Item Primary path entry Backup path entry Approval question
Selected flow Approved client and destination scope Same test scope Is the change bounded to the intended traffic?
PBF decision Intended rule, order and egress Remaining matching rules or no PBF match What actually happens when the primary rule is disabled?
Monitor Target, profile, action and disable setting Independent backup health check Which failure domain does each check cover?
Routing Intended virtual router and usable next hop Usable fallback route and adjacency Can the next hop really forward traffic?
NAT Expected translated source Expected translated source on backup Is the source valid and returnable through that ISP?
Security policy Intended allow rule and inspection Equivalent approved access Does the fallback preserve security controls?
Application Named transaction and expected result Reconnect behavior and recovery objective What counts as recovered for the user?
Recovery Restoration steps and observation period Rollback path and access method Can the operator recover without the failed circuit?

Recommended workflow:

  1. Save the approved baseline and identify a tested management path that does not depend solely on the circuit under test.
  2. Verify that the candidate policy is scoped to the intended flow and virtual system. Review dependencies and other pending changes before committing.
  3. Validate the backup route, next-hop reachability and NAT expectations while the primary path is healthy. Do not assume a configured default route is usable.
  4. Confirm monitor target, monitoring profile action, disable setting and remaining PBF rule order.
  5. Apply only the approved change through the normal commit workflow. Check the resulting running configuration and job outcome.
  6. Test one failure mode at a time, retain evidence, then restore and observe recovery before moving to the next test.

For broader configuration deployment problems, see the Panorama commit and push troubleshooting checklist. For selecting related network topics, use Start Here.

When symmetric return is relevant—and when it is not

Consider a different flow from outbound branch internet access: a client reaches a published server through ISP B, but the reply would otherwise leave through ISP A. Palo Alto describes symmetric return as forwarding the return traffic out through the interface on which the traffic originally arrived.[3] It is relevant to this return-path problem, not a generic fix for every dual-ISP failure.

The LIVEcommunity article also describes using a No PBF forwarding action when the client-to-server path should follow normal routing while the reverse path benefits from symmetric return.[3] Do not copy that choice into an unrelated outbound steering rule without tracing both directions.

For an approved read-only inspection, the vendor documents this operational command:[3]

show pbf return-mac all

This command is an unexecuted example here; no firewall output or lab result is being claimed. Compare any relevant entries with the actual flow and observed egress rather than treating an entry's existence as proof that the application is healthy.

The vendor article notes next-hop caveats and a same-subnet condition under which symmetric return is not used.[3] Review those restrictions for the actual topology and release before enabling the feature. If the next hop is unresolved or the observed MAC belongs to an unexpected intermediate device, investigate that topology first. Do not disable security inspection to compensate for an unexplained return path.

Failover acceptance checklist: test more than a cable pull

These are proposed tests for an isolated lab or an approved maintenance window. Do not interrupt a shared production uplink without an agreed impact boundary and recovery path. All result cells must be populated from your own observations.

Test Controlled condition Required observations Suggested pass criterion
Healthy baseline Both circuits usable Matching policy, egress, NAT and application transaction Selected flows use the intended path and complete the transaction
Local link failure Approved primary link interruption Link state, monitor, fresh session and old session New transactions recover within the agreed target; old-session outcome is recorded
Upstream failure Primary Ethernet remains up; approved upstream failure is simulated Probe reachability and actual application reachability Detection matches the written failure-domain design
Probe-only failure Controlled loss of the monitoring target, not the entire service Monitor transition, rule state and application impact False-failure consequences are understood and acceptable
Backup-path validation Primary unavailable Backup egress, NAT identity, replies and inspection Full application transaction succeeds without weakening policy
Inbound dual-ISP access Authorized external test via each ISP separately Request ingress and response egress Each published service returns over its intended path
Restoration Primary becomes reachable again Monitor recovery, path selection and application reconnects Recovery is stable and meets the agreed failback behavior
Rollback Restore the approved baseline Commit outcome, access and transaction Known-good behavior and management access are restored

Record monitor-down time, the start of the first successful fresh transaction, and the time the application owner confirms usable service. Those are different milestones. Avoid presenting a vendor timer setting as a measured recovery time.

If a connection still fails, return to the first point where evidence diverges: rule selection, monitor result, egress, source translation, return traffic or application response. If the Traffic log shows an aged-out session, use the NAT-aware packet capture checklist as a companion rather than treating the session end reason as the root cause.

Practical takeaways

  • Prove which policy and path the affected flow actually uses.
  • Review monitoring action, disable setting and remaining PBF rules as one decision.
  • Separate newly created connections from established sessions in every test.
  • Validate backup NAT, return traffic and the business transaction—not only the default route.
  • Apply symmetric return only to an identified reverse-path requirement.
  • Keep the completed worksheet and acceptance matrix with the change record so the next incident starts with evidence, not guesses.

Sources

  1. [1] https://docs.paloaltonetworks.com/pan-os/11-1/pan-os-admin/policy/policy-based-forwarding/pbf/path-monitoring-for-pbf
  2. [2] https://docs.paloaltonetworks.com/pan-os/11-1/pan-os-admin/policy/policy-based-forwarding/pbf
  3. [3] https://live.paloaltonetworks.com/t5/general-articles/policy-based-forwarding-symmetric-return-overview/ta-p/545067

Comments

0 Responses to "Palo Alto PBF Failover Not Working: Dual-ISP Troubleshooting Checklist"

Post a Comment

Popular Posts