This first report establishes the starting point for a recurring view of the AI experiment behind The Glowing Oracle Lab. The experiment is active, all six providers have returned stored responses, and the first two checkpoint days are complete. The evidence is still early, with review work outstanding.
Question
How consistently do six AI providers solve the same controlled quantitative task, state their assumptions, and follow instructions when the task is repeated over time?
Method
The first cohort began September 27, 2026. Its prompt asks for the expansion of a 120-meter steel pedestrian bridge from 10°C to 35°C, given a linear expansion coefficient of 12 × 10⁻⁶ per °C. The task requires a formula, calculation, answer in millimeters, and assumptions, without web browsing. Under those supplied conditions, the reference calculation is 36 millimeters.
OpenAI, Anthropic, Google, xAI, Mistral, and DeepSeek participate. Each checkpoint has fresh sessions at minute 0, 5, 15, and 30. The stored schedule includes Day 0, 1, 7, 14, 21, 30, 90, 180, and 365. Short repeats help distinguish ordinary sample variation from changes observed over longer intervals.
Updated observations
| Measure | Verified baseline |
|---|---|
| Live cohorts | 1 |
| Completed observation slots | 8, spanning Day 0 and Day 1 |
| Successful generation attempts / stored responses | 48 / 48 |
| Coverage | 8 responses from each of the six providers |
| Future observation slots | 28 pending; none overdue at this check |
| Most recent stored response | September 28, 2026, 6:25 AM Eastern |
The runner’s every-minute schedule is active, and the three most recent inspected cron invocations succeeded. That confirms scheduler activity; the stored response records provide the stronger evidence of completed generation work. This report checks execution metadata, not the correctness of every answer.
Change since the previous report
This is the first published snapshot. There is no previous report for a measured comparison. Future reports will show changes in response coverage, checkpoint completion, evaluation coverage, review obligations, and findings supported by inspected evidence.
Review status
There are 36 evaluation records: 34 completed and two marked as requiring human review. No human-review records are stored. The most recent evaluation timestamp is September 27, so Day 1 evaluation coverage needs to be checked against the intended evaluation scope. “Completed” describes the recorded evaluator status; it is not a claim that Ben has independently validated the finding.
Limitations
One prompt cannot represent broad model capability. This audit has not adjudicated all raw outputs or validated the evaluator’s scoring. It does not establish provider rankings, correctness rates, or a longitudinal drift finding. Agreement and stability must be examined separately from correctness. The daily multi-model Council remains unverified and is outside this report’s observed experiment.
Leadership implications
The useful operating lesson is already visible: a scheduled task, stored outputs, automated assessments, and completed human review are separate states. A team needs to see each of them. The research architecture should preserve those distinctions so progress cannot quietly substitute for evidence.
As the study matures, the practical question will be whether the behavior a workflow depends on remains acceptable—and whether the people operating it can detect and respond to meaningful changes.
Next checkpoint and follow-up
Day 7 begins Sunday, October 4 at 5:54 AM Eastern, with repeats at 5:59, 6:09, and 6:24 AM. The next report is scheduled for November 1, followed by a new report every 30 days. Outstanding work includes the two flagged human reviews and verification of Day 1 evaluation coverage. If a future reporting window contains no new observations, its report will state that rather than manufacture a trend.