Median Request Elapsed Time
Sol: July 10–September 4; 47 session tags. Astra: September 4–7; 74 tags. The timer includes preparation, streaming, and possible retries/backoff, but excludes subsequent tools. It is not end-to-end response or task time.
Astra / Sol · Quality, Speed And Usage
Astra can feel smarter and faster while consuming an allowance faster. The published rates support the cost concern. ForgeBot’s records need stricter interpretation before they can prove a quality or speed advantage.
Internal · Actual Retained Records
Read-only analysis found 3,524 directly model-tagged requests across retained Astra and Sol logs on this host, filtered to the Codex provider. This is neither Carl’s machine nor a team-wide census. Astra’s sample is recent; Sol’s spans older retained dates. [I1]
Sol: July 10–September 4; 47 session tags. Astra: September 4–7; 74 tags. The timer includes preparation, streaming, and possible retries/backoff, but excludes subsequent tools. It is not end-to-end response or task time.
Astra’s observed calls carry smaller contexts and outputs: mean output 241 tokens vs Sol’s 387. Different task mix, prompt size, dates, and runtime versions prevent a model-only comparison.
September 4 is the only shared retained log date. Sol has 94 calls in 2 session tags; Astra has 469 in 43. The 90th percentile reverses direction: Astra 13.2 s versus Sol 12.47 s. This is still not matched work; it shows why the pooled median must not become a universal speed claim.
Recorded cached-input token fraction, not request hit rate or quota discount. Cache suffix is absent on 85 Sol calls and 65 Astra calls; absent does not prove the provider explicitly reported zero. Cached input is already within prompt totals.
OpenAI reports benchmark gains, and Carl supplies a firsthand impression. The retrieved Slack evidence does not contain matched task-quality scores. A transition-period thread also records a response-ownership mistake, so “smarter” should not be confused with flawless agent behavior. [I2]
Carl’s screenshot reports a useful follow-up question. The exact Miguel exchange was not independently identified. OpenAI explicitly describes Astra asking focused questions when answers change outcomes. ForgeBot instructions also encourage material clarification, so the model is not the only possible cause. [9]
The retained medians are lower for Astra. Adam also reported a faster feel after a September 4 upgrade. Neither establishes total task speed; our sample is unbalanced and changes across the cutover. [I3]
Published message estimates support less Astra headroom, but Carl’s image contains no usage gauges. Work bursts can stress the shorter window while the weekly allowance remains comfortable. Do not report “2.5× quota drain”: the promotional credit ratio is not that measurement.
Sol can stretch remaining allowance. It is not documented as a fresh quota after exhaustion. A retrieved September 4 fallback event named a different provider, not Sol, and reported provider failure—not proven quota exhaustion. That historical event does not establish today’s policy. [I4]
Snapshot: 2026-09-07T23:19:51.301Z. Read-only SQLite and strict accounting-line extraction from retained logs. Daily sums were independently reconciled to model aggregates. Direct model-tagged log calls were used instead of attributing all session totals to a session’s latest model label.
All stored message token-count fields were null at capture; route rows are cumulative and some are migration-seeded. Tool-call metadata is not a unique execution ledger. Therefore no invented daily subscription usage, tools-per-success figure, or task-quality score is shown.
Slack research read relevant speed, transition, fallback, and clarification threads. Fireflies keyword discovery found no relevant meeting; no full meeting transcript was retrieved. No authenticated OpenAI billing dashboard was opened. Narrow public sources were preferred over unsourced social comparisons.
External · Vendor Evaluations, Not ForgeBot Tests
OpenAI’s Astra announcement directly describes better judgment about when to clarify. It also reports stronger outcomes and faster computer use. These are vendor results in particular harnesses and settings, not independent replications or guarantees for ForgeFX work. [9]
OpenAI reports Astra at approximately 9% lower estimated API cost per task in this comparison despite higher unit pricing.
Benchmark score, scale 0–100%. API-cost result is benchmark/configuration-specific—not purchased credits, subscription drain, or locally measured tools-per-task. Sol comparator is the API/Codex/Work version.
OpenAI describes about 47% less time per task with a higher score.
Simulated latency in a computer-use evaluation, not customer response time. Do not compare these minutes with Hermes per-request seconds above.
External · Published Provider Evidence
OpenAI publishes separate credit rates for uncached input, cached input, and output. Astra’s standard rate is higher in all three categories. These are purchased-credit rates—not an invoice for this research and not a measured conversion to the subscription usage meter. OpenAI explicitly says Sol’s credit promotion leaves included plan usage, five-hour limits, and weekly limits unchanged. The 2.5× promotional rate ratio must not be presented as a quota multiplier. [1] [10]
Category-specific scale starts at zero.
Category-specific scale starts at zero.
Each chart has its own scale. Within each category, Astra costs 2.5× Sol.
Source: OpenAI Codex pricing, retrieved September 7, 2026. Sol’s listed promotional rates run at least through November 21, 2026. Credit rates exclude other charges and contract differences. [1]
At the published standard rates, Astra must use 60% fewer cost-weighted tokens to match Sol’s token-credit cost for the same accepted task. With the same input/cache/output proportions, that means 60% fewer tokens.
A smaller reduction can still be worth paying for if Astra avoids expensive mistakes or saves staff time. Fewer tool calls alone do not establish either result.
For Plus and Standard Business, OpenAI estimates 5–45 local messages with Astra versus 10–100 with Sol per five-hour period. Tasks differ in context, reasoning, tools, retrieval, and cache use. These ranges are not hard message caps. [1]
Astra: 5–45
Sol: 10–100
Weekly limits may also apply. Cloud chats and local messages share an allowance. The usage dashboard is authoritative for the active account’s current limits and reset times.
Infographic · Why Both Observations Can Be True
Better planning can remove retries and tool calls. A useful clarification can prevent a wrong branch. These are hypotheses to test against accepted outcomes.
Astra’s published standard token-credit rates are 2.5× Sol’s. Context size, cache hits, output and reasoning can outweigh fewer calls.
Faster completion and parallel workers can pack more work into a short window. Other sessions using the same allowance add to consumption.
Quota used per hour = tasks per hour × average quota debit per task + other workload
Accounting relationship only. No task throughput or quota debit values have been invented. Five-hour and weekly meters have different windows and denominators; their percentages are not directly comparable.
External + Internal · Attribution
A messaging platform is the entry point. It does not establish which OpenAI subscription pays for inference.
The configured inference credential determines the account route. Delegated agents and auxiliary calls can add load; verify each route rather than assuming every call is separate.
Separate the bot’s allowance from the workspace’s pooled purchased credits and other employees’ individual allowances. Workspace policy and seat type matter.
OpenAI supports separate personal and work accounts. The account switcher keeps billing, history, memory, and workspaces separate; it does not merge subscriptions. Its help page currently says ChatGPT web only, not Codex desktop. [3]
Hermes supports same-provider credential pools, but that technical capability does not establish authorization for shared employee access or moving company context into a personal account. Keep work under the approved work identity; do not treat rotation as permission to bypass limits. [4]
OpenAI’s Business page currently advertises a Premium seat with 5× Standard usage and no five-hour limit. Published prices are $100/month billed annually or $125 billed monthly, versus Standard at $20/$25. [2]
The Codex pricing page still presents five-hour estimate tables and says Business ($100) uses Pro 5× estimates. These are different presentation surfaces, not proof of ForgeFX’s entitlement. Verify seat type, active workspace, and the live usage dashboard before buying or changing anything. [1]
Recommendation: Keep routine work on the approved lower-cost route and reserve Astra for tasks where its outcome or time savings justify the premium. Treat Sol fallback as a deliberate model-selection policy—not a promise that an exhausted shared allowance resets. No account, model, billing, or gateway settings were changed for this research.
Decision · What Would Settle The Question
This controlled comparison is proposed, not executed. Research used existing records and public documentation instead of spending additional quota on an unrequested benchmark campaign.
Evidence Register
Public research prioritizes official OpenAI and Hermes documentation. Internal research uses read-only local session/log records and authorized Slack evidence. Metrics are calculated in code. This is an observational research preview, not a randomized Astra-versus-Sol evaluation.
Human reports are not task-level acceptance scores. Provider benchmarks are not ForgeBot benchmarks. Token totals are not subscription quota debit. Cached input is a subset of input, not extra usage to add twice. Session model labels can hide mixed-model histories. Active requests, retained-history limits, missing counters, and unequal task mix constrain comparisons.
Local source aggregates and calculation scripts remain in the report’s artifact directory. No raw conversations, secrets, personal account identifiers, client material, or credential values are published here. Billing-dashboard access and coverage gaps are described in the Internal section.