Why “We Have Failover” Still Fails in Real Life
The gap between networking and business continuity
Most Manhattan offices can buy a second ISP and a 5G router. The failure happens later—when an outage hits and nobody can confirm which workflows still function, who’s responsible for validating them, or whether the backup path can support the *minimum* your business needs.
What an acceptance test actually proves
An acceptance test is a business-owned drill that intentionally drops the primary circuit and verifies a defined list of workflows, in order, with pass/fail criteria and named owners. It reduces finger‑pointing between ISP, MSP, landlord/building staff, and internal teams because the results are documented and repeatable.
What “Minimum‑Viable Connectivity” Means (App by App)
Stop aiming for “everything works” on backup
Your backup link (especially 5G) may not match your primary bandwidth, latency, or inbound capabilities. The goal isn’t perfection—it’s keeping the company operational in a controlled “degraded mode” until the primary circuit returns.
A practical matrix you can tailor in 30 minutes
Below is a starting point. Customize it to your environment, then assign an “owner” for each row—someone who can truthfully confirm the workflow works during the drill.
Minimum‑Viable Connectivity (MVC) Matrix (template)
- VoIP calling (inbound/outbound): Must place/receive calls; call quality acceptable; voicemail reachable. Owner: Office manager / front desk lead.
- E911 / location routing: Emergency calling routes correctly per your provider’s configuration and office address. Owner: Operations + IT.
- Microsoft 365 (Outlook/Teams/SharePoint): Send/receive email; join Teams call; open key SharePoint/OneDrive files. Owner: Team lead from each department.
- POS / payments (if applicable): Process a transaction; settle batch later if needed. Owner: Store/finance lead.
- RDP/VDI / remote access: Staff can reach desktops/VDI with MFA; performance is usable. Owner: IT + one power user.
- Line‑of‑business (LOB) app: Login + complete one “money path” workflow (e.g., create invoice, update ticket, submit order). Owner: App owner.
- Printing/scanning (if critical): Print to at least one device; scan to email or folder. Owner: Office manager.
- Guest Wi‑Fi: Optional; can remain off during failover. Owner: IT.
NYC Office Realities That Can Break Failover (Even If the Router Is Perfect)
Demarc access and who can touch what
In many buildings, the ISP demarcation room and riser closets aren’t freely accessible. If only the landlord or building engineer can grant entry, your “rapid repair” assumptions may be wrong.
IDF/riser permissions and change windows
Even simple moves—patching a different handoff, adding a cross‑connect, or relocating a modem—can require approvals and scheduled access. Your test plan should avoid steps that depend on last‑minute building coordination.
Provider lead times and dependency chains
Backup circuits can be delayed by building wiring availability, conduit capacity, or required paperwork. Your acceptance test should document what you *can* validate now and what’s pending, so the business knows what’s truly covered.
The 60‑Minute Acceptance Test Runbook (Quarterly Drill)
What you need before you start
Pick a time with low business risk, but run at least one drill during business hours each year so you’re not only testing “after‑hours conditions.” Ensure you have admin access to firewall/router, VoIP portal, and any MFA/SSO admin tools needed to troubleshoot.
Roles to assign (so it’s business-owned)
Name these roles in the runbook—even if one person temporarily wears two hats.
- Test captain (usually IT/MSP): Runs the clock, executes network steps, records results.
- Business verifier(s): Confirms app workflows using the MVC matrix.
- Comms lead: Notifies staff, sets expectations (“degraded mode”), and announces start/end.
- Building/landlord contact (as needed): On standby for access issues, not for routine steps.
- Primary ISP account info + support number
- Backup ISP/5G account info + support number
- Firewall/router admin credentials + out-of-band access method
- List of critical apps + owners (MVC matrix)
- Screenshot or export of current WAN/failover configuration
- A place to log results (shared doc + timestamped folder)
Step 1: Establish the baseline (10 minutes)
Confirm “normal mode” is healthy
Run quick checks while still on the primary circuit: can you place a VoIP call, access Microsoft 365, and run the key LOB workflow? This prevents confusing “existing issues” with failover issues.
Record the starting state
Capture WAN status, public IPs, and current routing/failover status (screenshots are fine). Note which staff are participating and what devices they’re using (laptop on Wi‑Fi vs. desktop on Ethernet matters).
Step 2: Force failover and validate MVC (35 minutes)
Trigger failover intentionally (don’t wait for it)
The safest method is usually disabling the primary WAN interface on the firewall/router rather than physically unplugging cables in shared closets. The goal is a controlled switch you can reverse quickly.
Start the clock and confirm the network state
Note the time failover was initiated and the time users regain connectivity. Record what the firewall reports (primary down, backup up) and what users experience (DNS, browsing, app logins).
Test in this order (most revenue-impact first)
Test each item as pass/fail with a short note, and stop troubleshooting deep issues until you’ve completed the whole list once.
- 1) VoIP + E911 readiness: Place outbound call, receive inbound call, verify voicemail access. Confirm E911 settings *exist and are correct* in the provider portal (you generally don’t “test” emergency calls casually).
- 2) Authentication/MFA: Ensure SSO and MFA prompts work for at least one user from each team.
- 3) Microsoft 365 core: Email send/receive, Teams call audio, open/share a key file.
- 4) LOB “money path”: Complete one real workflow end-to-end (use a test record if possible).
- 5) RDP/VDI / remote access: Connect, operate for 2–3 minutes, note latency.
- 6) POS/payments (if applicable): Run a controlled test transaction per vendor guidance.
Validate “hidden dependencies” that cause finger‑pointing later
If you use any of the following, include a simple check:
- IP allowlists for vendors/banks
- Site-to-site VPNs
- On‑prem servers that must be reached from outside
- DNS filtering/security tools that may behave differently on 5G
Step 3: Fail back cleanly and close the loop (15 minutes)
Restore the primary link and confirm reversion
Re-enable the primary WAN interface and verify the firewall returns to preferred routing. Confirm that key apps work again and that no users are “stuck” on the wrong path.
Document what changed (even if everything passed)
Log firmware updates, ISP modem swaps, AP moves, VLAN changes, or new SaaS apps since the last drill. This change log is what makes the next quarter’s test faster and prevents “we didn’t touch anything” confusion.
Create tickets for failures with clear ownership
For each failure, assign an owner and next action: MSP config, vendor escalation, internal training, or building coordination. Include evidence: timestamps, screenshots, and the exact symptom.
Your Scorecard: What a Pass Looks Like
Use simple thresholds so results are comparable quarter to quarter
Define pass/fail with tolerances that match your business.
- Failover time: Users regain basic internet within X minutes (set your target).
- VoIP: Calls connect reliably; acceptable audio on a short call.
- M365: Email and Teams meeting join succeeds for test users.
- LOB app: One key workflow completes without workaround.
- RDP/VDI: Session connects; usable interaction (note “usable” vs. “ideal”).
Track “degraded mode” decisions explicitly
If guest Wi‑Fi, large file sync, or video meetings are nonessential, write that down. During a real outage, clarity prevents staff from saturating the backup link and accidentally breaking the workflows that matter.

How Often to Run This (and When to Re-Test)
Quarterly drills plus “after change” micro-tests
Run the full 60‑minute drill quarterly. Also run a 10‑minute micro-test after meaningful changes: firewall replacement, VoIP vendor changes, new MFA policies, office moves, or adding a new SaaS that becomes business-critical.
Keep it survivable through staff turnover
Store the runbook and last scorecards in a shared location with restricted admin access. Ensure at least two people know where it is and can run it, even if your office manager or IT contact is out.
Key Takeaways
- A backup circuit isn’t resilience until you prove critical workflows work on it.
- Build an app-by-app minimum‑viable connectivity matrix with named business owners.
- Run a timed, quarterly acceptance test that forces failover and records pass/fail evidence.
- Include NYC office constraints (demarc/riser access, permissions, lead times) in your plan.
- Maintain a change log so drills stay relevant as your environment evolves.

Frequently Asked Questions
How disruptive is a 60‑minute failover acceptance test?
If planned, it’s usually manageable because you’re intentionally operating in “degraded mode” for a short window. The disruption is the point—you’re trading a controlled inconvenience for confidence during an unplanned outage.
Do we need dual wired ISPs *and* 5G?
Not always, but many SMBs use 5G as a tertiary path or a fast interim solution while waiting for a second wired install. Your acceptance test will reveal whether 5G can meet your minimum‑viable needs or only serve as emergency email access.
What if our VoIP provider says E911 is “handled automatically”?
You still want an explicit verification step: confirm the service address and any location routing are correct in the provider portal and documented internally. Treat E911 as a compliance-and-safety item, not a “we’ll assume it works” feature.
Will failover break vendors that whitelist our public IP?
It can. If a vendor only trusts your primary static IP, switching to backup may block access until the vendor allowlist is updated or you implement a design that preserves expected egress identity.
Who should own the acceptance test: IT or the business?
IT should run the network actions, but the business must own the definition of “minimum viable” and provide verifiers for each workflow. That shared ownership is what prevents finger‑pointing when it matters.
Take the Next Step
If you already have dual ISP or 5G backup (or you’re budgeting for it) and want proof it will hold up in a real Manhattan workday, Your Expert Tech can help you build a minimum‑viable connectivity matrix, run a 60‑minute acceptance test, and turn the results into a quarterly scorecard your team can repeat.
Contact us to schedule a failover readiness review and get a drill plan your office can actually execute—without relying on tribal knowledge.

