/Users/forgebot/forgeapps/artifacts/astra-research/report.html
RESEARCH PREVIEW
September 7, 2026 · For Adam and Carl

Astra / Sol · Quality, Speed And Usage

Efficiency Is Not Quota.
Both Observations Can Be True.

Astra can feel smarter and faster while consuming an allowance faster. The published rates support the cost concern. ForgeBot’s records need stricter interpretation before they can prove a quality or speed advantage.

2.5×Astra vs Sol standard token-credit rates. Published, not measured quota drain. [1]
60%Cost-weighted-token reduction needed to break even. Derived at current standard rates.
Not Yet ProvenForgeBot’s causal quality/speed gain and change in team allowance drain.

Internal · Actual Retained Records

The Speed Signal Is Encouraging. The Comparison Is Uneven.

Read-only analysis found 3,524 directly model-tagged requests across retained Astra and Sol logs on this host, filtered to the Codex provider. This is neither Carl’s machine nor a team-wide census. Astra’s sample is recent; Sol’s spans older retained dates. [I1]

Observed · Not causal

Median Request Elapsed Time

Sol · 2,564 calls9 s
Astra · 960 calls5.8 s

Sol: July 10–September 4; 47 session tags. Astra: September 4–7; 74 tags. The timer includes preparation, streaming, and possible retries/backoff, but excludes subsequent tools. It is not end-to-end response or task time.

Confounder · Different workloads

Mean Prompt Tokens Per Call

Sol · 2,564 calls132,904.71 tokens
Astra · 960 calls72,656.38 tokens

Astra’s observed calls carry smaller contexts and outputs: mean output 241 tokens vs Sol’s 387. Different task mix, prompt size, dates, and runtime versions prevent a model-only comparison.

Same-Day Check: A Smaller Gap

Sol · 94 calls6.25 s
Astra · 469 calls5.5 s

September 4 is the only shared retained log date. Sol has 94 calls in 2 session tags; Astra has 469 in 43. The 90th percentile reverses direction: Astra 13.2 s versus Sol 12.47 s. This is still not matched work; it shows why the pooled median must not become a universal speed claim.

Most Prompt Tokens Were Cached

Sol · 2,564 calls94.8%
Astra · 960 calls92.93%

Recorded cached-input token fraction, not request hit rate or quota discount. Cache suffix is absent on 85 Sol calls and 65 Astra calls; absent does not prove the provider explicitly reported zero. Cached input is already within prompt totals.

Recent Astra Activity Is Bursty

Sep 04 · 469 calls29.41M tokens
Sep 05 · 15 calls0.45M tokens
Sep 06 · 125 calls11.38M tokens
Sep 07 · 351 calls28.51M tokens
0M10M20M30M29.41MSep 04n=469 calls0.45MSep 05n=15 calls11.38MSep 06n=125 calls28.51MSep 07n=351 calls
Full prompt tokens, including cached context repeatedly supplied to calls. Retained log-local dates (historical timezone not established); September 7 is partial at capture. Gaps and retention mean these are observed logged totals, not complete daily or team consumption. [I1]
Team usage remains unresolved. The saved default is Astra through Codex OAuth, with low reasoning. Retained route records include Slack, CLI, desktop, scheduled work and subagents; auxiliary activity also uses the Codex provider. That does not identify the subscription owner, prove a common credential, or establish the workspace’s paid-credit drain. No authoritative before/after allowance or billing series was obtained.

Answering Carl’s Topics

“It seems better.”

Supported direction · Not locally scored

OpenAI reports benchmark gains, and Carl supplies a firsthand impression. The retrieved Slack evidence does not contain matched task-quality scores. A transition-period thread also records a response-ownership mistake, so “smarter” should not be confused with flawless agent behavior. [I2]

“It asked Miguel.”

Anecdote + documented behavior

Carl’s screenshot reports a useful follow-up question. The exact Miguel exchange was not independently identified. OpenAI explicitly describes Astra asking focused questions when answers change outcomes. ForgeBot instructions also encourage material clarification, so the model is not the only possible cause. [9]

“It is faster.”

Observed request-time signal

The retained medians are lower for Astra. Adam also reported a faster feel after a September 4 upgrade. Neither establishes total task speed; our sample is unbalanced and changes across the cutover. [I3]

“Five hours drains.”

Plausible · Actual drain unmeasured

Published message estimates support less Astra headroom, but Carl’s image contains no usage gauges. Work bursts can stress the shorter window while the weekly allowance remains comfortable. Do not report “2.5× quota drain”: the promotional credit ratio is not that measurement.

“Fallback to Sol.”

Proposed strategy · Not verified config

Sol can stretch remaining allowance. It is not documented as a fresh quota after exhaustion. A retrieved September 4 fallback event named a different provider, not Sol, and reported provider failure—not proven quota exhaustion. That historical event does not establish today’s policy. [I4]

Internal Collection Limits And Data Download

Snapshot: 2026-09-07T23:19:51.301Z. Read-only SQLite and strict accounting-line extraction from retained logs. Daily sums were independently reconciled to model aggregates. Direct model-tagged log calls were used instead of attributing all session totals to a session’s latest model label.

All stored message token-count fields were null at capture; route rows are cumulative and some are migration-seeded. Tool-call metadata is not a unique execution ledger. Therefore no invented daily subscription usage, tools-per-success figure, or task-quality score is shown.

Slack research read relevant speed, transition, fallback, and clarification threads. Fireflies keyword discovery found no relevant meeting; no full meeting transcript was retrieved. No authenticated OpenAI billing dashboard was opened. Narrow public sources were preferred over unsourced social comparisons.

External · Vendor Evaluations, Not ForgeBot Tests

Carl’s Quality Impression Has A Public Basis

OpenAI’s Astra announcement directly describes better judgment about when to clarify. It also reports stronger outcomes and faster computer use. These are vendor results in particular harnesses and settings, not independent replications or guarantees for ForgeFX work. [9]

Terminal-Bench 4.0

Astra57.9%
Sol37.3%

OpenAI reports Astra at approximately 9% lower estimated API cost per task in this comparison despite higher unit pricing.

Benchmark score, scale 0–100%. API-cost result is benchmark/configuration-specific—not purchased credits, subscription drain, or locally measured tools-per-task. Sol comparator is the API/Codex/Work version.

OSWorld 2.0 Latency Simulation

Astra · score 72.6%~40 min/task
Sol · score 65.7%~75 min/task

OpenAI describes about 47% less time per task with a higher score.

Simulated latency in a computer-use evaluation, not customer response time. Do not compare these minutes with Hermes per-request seconds above.

The harness matters: OpenAI’s separate 1.9× faster Mind2Web claim combines Astra with an updated Codex computer-use harness. ForgeBot runs Hermes; using the same model does not establish that Hermes inherits Codex’s harness gains. [9]

External · Published Provider Evidence

The Price Of Each Token Matters

OpenAI publishes separate credit rates for uncached input, cached input, and output. Astra’s standard rate is higher in all three categories. These are purchased-credit rates—not an invoice for this research and not a measured conversion to the subscription usage meter. OpenAI explicitly says Sol’s credit promotion leaves included plan usage, five-hour limits, and weekly limits unchanged. The 2.5× promotional rate ratio must not be presented as a quota multiplier. [1] [10]

Uncached Input

GPT-6 Astra250 credits / 1M
GPT-5.6 Sol100 credits / 1M

Category-specific scale starts at zero.

Cached Input

GPT-6 Astra25 credits / 1M
GPT-5.6 Sol10 credits / 1M

Category-specific scale starts at zero.

Output

GPT-6 Astra1,250 credits / 1M
GPT-5.6 Sol500 credits / 1M

Each chart has its own scale. Within each category, Astra costs 2.5× Sol.

Source: OpenAI Codex pricing, retrieved September 7, 2026. Sol’s listed promotional rates run at least through November 21, 2026. Credit rates exclude other charges and contract differences. [1]

Derived analysis · Not a benchmark

The Break-Even Is A Large Reduction

At the published standard rates, Astra must use 60% fewer cost-weighted tokens to match Sol’s token-credit cost for the same accepted task. With the same input/cache/output proportions, that means 60% fewer tokens.

A smaller reduction can still be worth paying for if Astra avoids expensive mistakes or saves staff time. Fewer tool calls alone do not establish either result.

0% fewer weighted tokens2.5× Sol credits
20% fewer weighted tokens2× Sol credits
40% fewer weighted tokens1.5× Sol credits
60% fewer weighted tokens1× Sol credits
80% fewer weighted tokens0.5× Sol credits
0×1×2×2.5×Sol cost = 1×0%20%40%60%80%Reduction in cost-weighted tokens
Derived sensitivity, not observed efficiency. Relative credits = 2.5 × remaining cost-weighted tokens. Cache/output mix is held constant for a raw-token interpretation.
Provider estimates · Not task capacity

Five-Hour Message Ranges

For Plus and Standard Business, OpenAI estimates 5–45 local messages with Astra versus 10–100 with Sol per five-hour period. Tasks differ in context, reasoning, tools, retrieval, and cache use. These ranges are not hard message caps. [1]

Astra: 5–45

Sol: 10–100

Astra: 5–45Sol: 10–1000255075100Estimated local messages / five hours
Range endpoints, not confidence intervals. Model/task mix is not controlled.

Weekly limits may also apply. Cloud chats and local messages share an allowance. The usage dashboard is authoritative for the active account’s current limits and reset times.

Fast is not the same as Fast mode. Carl’s perceived speed does not establish that paid Fast mode is enabled. OpenAI lists a 2.5× multiplier for Astra Fast mode: that is 6.25× Sol’s standard token-credit rate and requires an 84% cost-weighted-token reduction to break even. This is a conditional calculation, not a claim about Carl’s settings. [1]

Infographic · Why Both Observations Can Be True

A Faster Agent Can Empty A Window Faster

01 / TASK EFFICIENCYLess Work Per Result

Better planning can remove retries and tool calls. A useful clarification can prevent a wrong branch. These are hypotheses to test against accepted outcomes.

02 / RESOURCE RATEMore Credits Per Token

Astra’s published standard token-credit rates are 2.5× Sol’s. Context size, cache hits, output and reasoning can outweigh fewer calls.

03 / WORKLOAD VOLUMEMore Work Per Hour

Faster completion and parallel workers can pack more work into a short window. Other sessions using the same allowance add to consumption.

Quota used per hour = tasks per hour × average quota debit per task + other workload

Accounting relationship only. No task throughput or quota debit values have been invented. Five-hour and weekly meters have different windows and denominators; their percentages are not directly comparable.

External + Internal · Attribution

Count The Credential, Not The Chat Room

REQUEST SOURCESSlack / Teams / CLI / Jobs

A messaging platform is the entry point. It does not establish which OpenAI subscription pays for inference.

EXECUTION ROUTEForgeBot → Model Provider

The configured inference credential determines the account route. Delegated agents and auxiliary calls can add load; verify each route rather than assuming every call is separate.

BILLING BOUNDARYSeat Allowance / Credits

Separate the bot’s allowance from the workspace’s pooled purchased credits and other employees’ individual allowances. Workspace policy and seat type matter.

Personal + Work Accounts

OpenAI supports separate personal and work accounts. The account switcher keeps billing, history, memory, and workspaces separate; it does not merge subscriptions. Its help page currently says ChatGPT web only, not Codex desktop. [3]

Hermes supports same-provider credential pools, but that technical capability does not establish authorization for shared employee access or moving company context into a personal account. Keep work under the approved work identity; do not treat rotation as permission to bypass limits. [4]

Check The Actual Seat First

OpenAI’s Business page currently advertises a Premium seat with 5× Standard usage and no five-hour limit. Published prices are $100/month billed annually or $125 billed monthly, versus Standard at $20/$25. [2]

The Codex pricing page still presents five-hour estimate tables and says Business ($100) uses Pro 5× estimates. These are different presentation surfaces, not proof of ForgeFX’s entitlement. Verify seat type, active workspace, and the live usage dashboard before buying or changing anything. [1]

Recommendation: Keep routine work on the approved lower-cost route and reserve Astra for tasks where its outcome or time savings justify the premium. Treat Sol fallback as a deliberate model-selection policy—not a promise that an exhausted shared allowance resets. No account, model, billing, or gateway settings were changed for this research.

Decision · What Would Settle The Question

Measure Accepted Work, Not Enthusiasm

Quality And Clarification

  • Use the same representative tasks and acceptance checks on both models; review answers without model labels.
  • Score correct completion, rework, and avoidable mistakes. A question is useful only when its answer changes the action or risk.
  • Separate model behavior from instructions: ForgeBot already has explicit guidance to ask when missing context materially changes the task. OpenAI’s Model Spec also endorses balancing clarification cost against a wrong assumption. [5]

Speed And Allowance

  • Hold account/plan, reasoning effort, tools, and task scope constant; randomize model order and avoid unrelated concurrent traffic.
  • Record active end-to-end time, model calls, tool calls, retries, and input/cache/output tokens per accepted task.
  • Capture provider five-hour and weekly percentages and reset times before/after each run. Exclude runs spanning a reset.
  • Compare paired results and uncertainty. Session lifespan is not response latency; a final reply is not proof of acceptance.

This controlled comparison is proposed, not executed. Research used existing records and public documentation instead of spending additional quota on an unrequested benchmark campaign.

Evidence Register

Sources And Boundaries

  1. OpenAI Codex pricing · Primary · current pricing
    Astra/Sol token-credit rates, Fast multiplier, local-message estimates, weekly limits, and Sol promotional period. Retrieved September 7, 2026.
  2. OpenAI Business pricing · Primary · current plans
    Standard/Premium pricing and Premium no-five-hour-limit statement. Confirm actual seat entitlement in the workspace.
  3. Use multiple accounts with account switching · Primary · account policy
    Personal/work separation, independent billing, and supported account-switching surface.
  4. Hermes credential pools · Primary · harness documentation
    Same-provider rotation, delegation sharing, and cache implications. Technical capability does not establish account-policy permission.
  5. OpenAI Model Spec — clarification · Primary · behavior specification
    Desired behavior, not empirical proof that one model asks better questions.
  6. Flexible pricing for Business, Enterprise and Edu · Primary · billing boundaries
    Business per-seat limits and optional shared purchased credits differ from Enterprise/Edu contracted pooled credits.
  7. Hermes fallback providers · Primary · harness documentation
    Fallback behavior, auxiliary routing, and model/account cache resets. Current local implementation/configuration may differ.
  8. OpenAI workspace analytics · Primary · measurement surfaces
    Analytics activity reports, billing/usage controls, and auditable records are separate surfaces. Availability is plan/role dependent.
  9. GPT-6 Astra announcement · Primary · vendor evaluation
    Clarification behavior, Terminal-Bench 4.0, OSWorld latency simulation, and combined model/harness Mind2Web claim. Not independently replicated here.
  10. Business / Enterprise credit-based rate card · Primary · scoped credit pricing
    Confirms rates and states Sol promotion applies to purchased credits; included plan usage, five-hour and weekly limits are unchanged.
  11. I1 · Local read-only telemetry collection · Internal · measured records
    Snapshot 2026-09-07T23:19:51.301Z. Retained agent accounting logs plus schema/instrumentation audit; chart data downloadable above. Detailed local methodology: telemetry.md in the artifact directory.
  12. I2 · Transition-period response-ownership correction · Internal · observed interaction
    September 4. Minimal paraphrase only; thread spans provider changes, so do not assign all behavior to Astra.
  13. I3 · Upgrade report and speed impression · Internal · human report
    September 4. Deployment and perceived speed evidence, not a timed model comparison.
  14. I4 · Historical fallback event · Internal · operational notice
    September 4. Provider-failure notice, not a quota diagnosis; not proof of current fallback settings.
Methodology And Coverage

Public research prioritizes official OpenAI and Hermes documentation. Internal research uses read-only local session/log records and authorized Slack evidence. Metrics are calculated in code. This is an observational research preview, not a randomized Astra-versus-Sol evaluation.

Human reports are not task-level acceptance scores. Provider benchmarks are not ForgeBot benchmarks. Token totals are not subscription quota debit. Cached input is a subset of input, not extra usage to add twice. Session model labels can hide mixed-model histories. Active requests, retained-history limits, missing counters, and unequal task mix constrain comparisons.

Local source aggregates and calculation scripts remain in the report’s artifact directory. No raw conversations, secrets, personal account identifiers, client material, or credential values are published here. Billing-dashboard access and coverage gaps are described in the Internal section.