Incident review · August 24–25, 2026

ForgeBot Model Stall

What actually broke, what fixed it, and which plausible-looking clues sent the investigation sideways.

120sPrimary first-byte timeout
1 → 1Retry policy after repair
5.2sRecovered model turn
14/14Regression checks passing

Executive Answer

The root problem was not Slack. The long-lived Hermes gateway was running multiple openai-codex turns through the same subscription/OAuth transport. Those concurrent streams could wedge: the provider accepted the connection but produced no first byte or no SSE events. Slack had already delivered the message; the model turn stalled inside the gateway.

The effective fix was to serialize Codex gateway turns, then restart the stock gateway to clear poisoned in-flight state. The gateway now applies a process-wide lock only to openai-codex, preserving parallelism for other providers. Retries were reduced from three to one, and subscription-backed Grok was configured as fallback resilience—not as the root repair.
Root cause

Shared Codex subscription transport under concurrency

Logs repeatedly show connections accepted by the Codex endpoint but no stream bytes for 120 seconds, or a first byte followed by no SSE events for 60 seconds. Clean one-shot model probes worked while Slack gateway turns stalled. That isolates the failure to the long-lived, concurrent gateway execution path rather than the model account or Slack transport.

Solution

Provider-scoped serialization plus a clean restart

A process-wide lock now wraps agent.run_conversation only when the active provider is openai-codex. The restart cleared abandoned in-flight connections. The interrupted exact-response request resumed and completed in 5.2 seconds; later fresh Slack turns completed normally across several channels.

Evidence Timeline

HSTObserved factMeaning
05:36–06:10Gateway reported no stored Codex credentials in early runs.A real earlier fault, but not the later stall after OAuth was restored.
06:16 onwardRepeated 120-second no-first-byte and 60-second no-SSE timeouts; some turns lasted 7–18 minutes.The request entered Hermes and reached the Codex endpoint, but streaming wedged.
14:29Provider-scoped Codex serialization was added to the installed gateway runtime.Repairs the unsafe execution shape; a restart is required to load it.
15:13A direct xai-oauth / grok-4.6 probe returned the exact expected token.Proved fallback entitlement at that moment; did not prove automatic fallback inside the gateway.
21:17–21:20STANDARD_MODEL_OK test timed out in the gateway; fallback was unavailable to that process.Slack ingress worked. Primary streaming still stalled, and fallback visibility was a separate resilience gap.
21:49–21:54STANDARD_MODEL_FIXED test stalled until the gateway was restarted.The old process still held bad in-flight state.
21:54:47–21:54:57Fresh stock gateway started, auto-resumed one interrupted session, and returned the exact 20-character answer in 5.2 seconds.Strong confirmation that process/transport state—not Slack routing—was the active blocker.
02:11 onwardFresh Slack requests completed in 3.2–40.6 seconds across multiple channels.Recovery persisted beyond the synthetic auto-resume turn.

Biggest Decoys / Red Herrings

Decoy #1

Slack credentials and Socket Mode

Slack authenticated as ForgeBot, received Adam’s messages, threaded warnings correctly, and could send replies. Credential cleanup and Doppler mapping were worthwhile, but they did not explain provider connections that accepted requests and then emitted no stream data.

Decoy #2

“The model is down”

Direct one-shot Codex probes succeeded while the persistent gateway path failed. Model catalog presence also proved only discoverability, not a healthy stream. The failure was execution-path-specific.

Decoy #3

Grok billing / entitlement messages

The fallback produced both “not configured” and spending-limit-style errors during the incident, even though a direct subscription probe succeeded earlier. That was a genuine fallback problem, but fallback failure did not cause the primary Codex stream to wedge.

Decoy #4

Context size and compression warnings

Failures appeared around 16.5K–34.6K estimated tokens—well below the configured model context—and also hit trivial prompts. “Context engine not found” was noisy but not causal.

Decoy #5

Successful outbound probes

A bot-authored post proves only that Slack can accept a message. It cannot prove human ingress, model execution, correct threading, latency, or duplicate safety. Earlier health checks overclaimed success from component checks.

Decoy #6

The first exact “fixed” response by itself

STANDARD_MODEL_FIXED arrived through restart auto-resume from an empty synthetic inbound event. It was useful evidence, but not a complete acceptance test. The later fresh human messages are what confirmed normal operation.

What Changed

Primary repair

  • Serialize openai-codex gateway turns with one process-wide lock.
  • Restart the gateway so the new runtime and a clean transport are active.

Resilience hardening

  • Use the stock ai.hermes.gateway supervisor only.
  • Keep Slack and Teams under one writer process.
  • Reduce provider retries from three to one.
  • Configure xai-oauth / grok-4.6 as the first fallback.
  • Load platform credentials through Doppler’s command secret source.

Verification hardening

The acceptance contract now requires one human-authored first-token mention, one exact threaded response in no more than 60 seconds, and a 120-second duplicate-observation window. Four regression suites currently pass all 14 checks: serialization, fallback configuration/entitlement classification, stock gateway control-plane invariants, and human Slack round-trip evaluation.

Source Coverage

Internal · decisive

  • Slack source threads and timestamps.
  • Hermes gateway and error logs.
  • Installed gateway runtime diff.
  • ForgeApps config and regression tests.
  • Session history containing direct provider probes.

External · supporting only

Hermes’ official configuration/provider guidance establishes the intended provider, OAuth, gateway, retry, and restart mechanisms. Public web retrieval was unavailable during this research pass, so no external source is used to claim the incident’s root cause. The root conclusion rests on first-party runtime evidence.

Confidence and Caveat

Confidence: high (0.88). The evidence strongly isolates the failure to the gateway’s long-lived Codex streaming path and shows recovery after serialization plus restart. The exact low-level failure inside the upstream Codex transport was not packet-captured, so “concurrent streams sharing one OAuth credential wedge the transport” remains a well-supported engineering diagnosis rather than a provider-confirmed postmortem. The fallback entitlement state also changed over time and should be treated separately.