In 20 years, you will be more dissapointed by what you didn't do than by what you did.

Prisma SD-WAN LTE Failover Not Working: Brownout Troubleshooting Matrix

A branch application becomes unusable, but the cellular circuit stays idle. Before changing timers or forcing traffic onto LTE, answer a narrower question: is cellular meant to rescue poor application performance, or only loss of the permitted primary paths? Those are different acceptance tests.

Palo Alto Networks' published SaaS example deliberately keeps metered 5G in the Layer 3 Failure Paths list: that example uses cellular when all active paths are down, not merely degraded.[2] This guide turns that distinction into an original troubleshooting matrix, policy worksheet and controlled failover test plan.

Scope: Prisma SD-WAN ION, not PAN-OS SD-WAN

This article concerns Prisma SD-WAN ION path and performance policy. The cited performance-policy guidance applies to ION software 6.3.1 and higher; the cellular-cost guidance includes older release-specific instructions.[1][3] Confirm your ION version and management interface before using any setting names below. Do not translate this worksheet directly into PAN-OS SD-WAN configuration.

The workflow, topology and acceptance criteria below are engineering recommendations, not measured product benchmarks. All test scenarios are hypothetical and must be adapted to an approved maintenance window. No customer topology, device address or private course diagram is reproduced.

The first decision: brownout or outage?

Use brownout here to mean that connectivity remains available but the application is slow or unreliable. Use outage to mean that the relevant allowed connectivity has failed. Record the product's observed path state separately from the user's description: “the internet is down” is a symptom report, not proof of an L3 failure.

In the vendor's example, two direct internet paths are active, there are no backup paths, and metered 5G is reserved for Layer 3 failure.[2] With that intent, a slow primary circuit does not automatically justify cellular use.[2] Your design may require a different commercial trade-off, but that must be explicit rather than inferred from the presence of a modem.

A useful requirement statement is: “Keep approved business applications usable during an uplink outage; use cellular during degradation only if the business has approved its cost and the deployed policy supports the intended behavior.” Write the requirement before adjusting the rule.

Original decision diagram

Decision diagram separating a branch application brownout from loss of primary paths
Source: original Network freak diagram; generic diagnostic workflow based on the cited public guidance. No private topology or customer data.

LTE failover troubleshooting matrix

Treat the following rows as diagnostic hypotheses, not automatic configuration fixes. Preserve evidence before changing the relevant policy.

Symptom Evidence to collect First decision or check Acceptance evidence
SaaS is slow; LTE is idle Application, matched path rule, current path state, latency/loss evidence Is LTE only an L3 failure path? Compare observed behavior with the documented outage-only example.[2] Business confirms whether degradation should or should not authorize cellular
An SLA alarm appears but sessions stay on the same circuit Matched performance rule, action selection, flow timestamps Compare actions with the vendor example, which explicitly selects Raise Alarms and Move Flows.[2] New and existing sessions are observed separately; alarm alone is not the acceptance result
An application-specific rule appears ineffective Site binding, stack order, application identification, first matching rule More-specific performance rules belong before less-specific rules; empty match fields mean match-all.[1] Observed flow matches the intended rule and path type
One wired circuit is removed but LTE stays idle Status of the other wired circuit and application transaction Is another eligible wired path still working? Do not use this test to claim all-path failure Application passes on the remaining approved path
All intended wired paths are unavailable; new sessions fail Cellular registration and reachability, assigned circuit category, effective path policy, endpoint access Prove cellular transport and intended rule/category separately; review the actual site binding.[3] A new approved transaction succeeds over the expected cellular path
LTE carries traffic but the application still fails DNS response, selected egress, destination reachability, authentication result Separate path selection from application acceptance DNS, connection setup and an authenticated transaction all pass
Failover passes; recovery is unstable Repeated tests, path history, rule changes, application errors Restore baseline and change only one variable per test Agreed stable observation period, with no unexplained path oscillation

Build a policy worksheet before touching production

Create one row per important application rather than one row labelled “internet.” Use an application owner to define success. A payment workflow, a remote administration session and a bulk backup may have different fallback requirements.

Worksheet field What to record
Application and owner Business workflow, test operator and escalation contact
Access model Direct internet, private application over overlay, or approved security-service path
Effective policy Site binding, stack, rule order, application match and path-type match
Normal eligible paths Circuit categories actually assigned at this branch
Degradation action Alert only, approved alternate-path behavior, or explicit outage-only cellular intent
Cellular authorization Which applications may consume metered capacity and who accepts the cost
Measurement Chosen metric, observation window, healthy baseline and failure evidence
Success test Specific transaction, observation method and agreed recovery target
Rollback Original policy reference, restore operator, management access and stop condition

Performance Policy has explicit rule ordering, and the documentation says the older Advanced-menu LQM/APT configuration is no longer used from 6.3.1; those rules must instead be configured in a performance policy set applied to the site.[1] Therefore, an old screenshot of configured thresholds is not enough: identify what is effective on the branch now.

The vendor SaaS example combines link-quality metrics with real-user application metrics and applies its rule to direct public paths.[2] Do not copy its numerical SLA thresholds into every branch. Capture a healthy baseline and use the application's agreed tolerance; otherwise the test may prove only that you chose an unsuitable threshold.

Separate three layers of evidence

1. The transport exists

First prove that the cellular service can reach the intended next service or destination under the deployed design. Record its registration state, circuit assignment and observed reachability using your supported diagnostics. A controller icon is not a substitute for the application test you intend to pass.

Check the cost-minimization settings carefully. Palo Alto Networks' cellular guide includes options to disable bandwidth and link-quality monitoring, controller connections and application reachability probes.[3] Those settings should trigger a review of what evidence you expect to see, not an immediate decision to enable everything.

2. Policy permits the intended path

Record the actual site-bound policy, not just the shared template. The cellular guide explicitly calls for reviewing path-policy stack bindings at the site.[3] Compare the rule's circuit category with the circuit configured at this branch. Keep the change small enough that an observed path change can be attributed to it.

The cellular guide warns that disabling application reachability probes for a circuit category should only be done when that category is referenced as an L3 failure path; doing so means the ION does not react to layer-7 failures for application flows on that circuit.[3] Treat cost optimization and application-failure detection as a design review together, not as independent checkboxes.

3. The application really works

Open a new test transaction after the failure event and also observe any session established before it. Record both outcomes. Do not promise uninterrupted authenticated sessions solely because a new connection succeeds.

Use application-site details, link-quality information and the flow browser to correlate the test: these are the monitoring views identified in the official SaaS example.[2] For the application owner, retain a simple result: what succeeded, what failed, which path was observed, and whether recovery met the agreed target.

Controlled brownout and outage test plan

Safety gate: retain independent management access, confirm the cellular allowance, notify application owners, and assign a rollback operator. Do not impair a shared carrier service or disable both wired paths outside an approved window. Use a lab or an isolated test branch for impairment injection.

Test Controlled stimulus What to prove Stop or rollback condition
Healthy baseline No induced fault Matched rules, normal paths and application transaction are understood Existing unexplained errors make the baseline invalid
Application brownout Approved impairment of the selected test path, while transport remains up Intended degradation response and explicit cellular yes/no requirement Unrelated applications or users are affected
Single wired-path loss Remove one approved test uplink Alternate eligible path and real transaction success Management access or unapproved services are lost
All intended wired-path loss Isolate relevant primary paths in the approved test setup New transaction over authorized cellular path No safe recovery access or unacceptable cellular use
Long-lived session observation Keep a permitted test session open across the selected event Distinguish session interruption from new-session recovery Sensitive production workflow is affected
Restoration Restore each original path deliberately Application recovery and stable subsequent path selection Repeated instability or unexplained policy behavior

Timestamp stimulus, detected path state, first successful new transaction and any existing-session interruption separately. This avoids conflating detection time with end-to-end application recovery. These are measurement fields to fill in, not benchmark results from this article.

Do not change policy order, probe behavior and impairment thresholds during the same test. Export or capture the approved baseline, record the smallest change, run the test, then retain or revert it based on evidence.

Do not enable FEC as a substitute for diagnosis

Forward error correction and packet duplication are separate capabilities from deciding whether cellular is eligible. The performance-policy documentation describes always-on and adaptive modes; adaptive behavior is tied to a loss threshold on a Prisma SD-WAN VPN path, and platform limits apply.[1] That is not a blanket fix for every direct SaaS brownout.

Before considering either feature, identify the actual path type, the measured impairment and the supported platform behavior. Keep this troubleshooting exercise focused on the question that started it: was the flow allowed and expected to use LTE under the failure you created?

Copyable acceptance checklist

  • [ ] ION software version and applicable documentation are recorded.
  • [ ] Each critical application has an owner and repeatable transaction.
  • [ ] Site-bound path and performance policies are identified.
  • [ ] Effective rule order and actual application/path matches are captured.
  • [ ] Cellular intent is explicit: outage-only or an approved degradation design.
  • [ ] Monitoring and reachability-probe settings match that intent.
  • [ ] Cellular transport is verified independently of path selection.
  • [ ] Brownout, single-path outage and all-primary-path outage are distinct tests.
  • [ ] New connections and existing sessions have separate results.
  • [ ] Application access and security requirements pass on fallback.
  • [ ] Restore behavior, cost exposure and rollback are approved.
  • [ ] Evidence includes timestamps and observed paths, not just green status icons.

Practical takeaways

Start with the permitted behavior, then prove the transport, the effective policy and the application result in that order. An idle LTE link during a brownout can be consistent with the documented outage-only design.[2] A successful failover test must demonstrate the business transaction, not merely a changed path indicator.

For the earlier preparation workflow, see Prisma SD-WAN branch migration checklist. For adjacent troubleshooting, use the IPsec tunnel-up/no-traffic matrix. Browse Start Here and the networking resources hub for related operational checklists.

Sources

  1. [1] Palo Alto Networks: best practices and recommendations
  2. [2] Palo Alto Networks: use case 1 protecting a business critical saas application
  3. [3] Palo Alto Networks: minimize metered lte usage

Comments

0 Responses to "Prisma SD-WAN LTE Failover Not Working: Brownout Troubleshooting Matrix"

Post a Comment

Popular Posts