Research / Experiment Report

AI Experiment Report — October 2, 2026

Ben Tartaglia · Baseline report · Data checked October 2, 2026

This first report establishes the starting point for a recurring view of the AI experiment behind The Glowing Oracle Lab. The experiment is active, all six providers have returned stored responses, and the first two checkpoint days are complete. The evidence is still early, with review work outstanding.

Question

How consistently do six AI providers solve the same controlled quantitative task, state their assumptions, and follow instructions when the task is repeated over time?

Method

The first cohort began September 27, 2026. Its prompt asks for the expansion of a 120-meter steel pedestrian bridge from 10°C to 35°C, given a linear expansion coefficient of 12 × 10⁻⁶ per °C. The task requires a formula, calculation, answer in millimeters, and assumptions, without web browsing. Under those supplied conditions, the reference calculation is 36 millimeters.

OpenAI, Anthropic, Google, xAI, Mistral, and DeepSeek participate. Each checkpoint has fresh sessions at minute 0, 5, 15, and 30. The stored schedule includes Day 0, 1, 7, 14, 21, 30, 90, 180, and 365. Short repeats help distinguish ordinary sample variation from changes observed over longer intervals.

Updated observations

MeasureVerified baseline
Live cohorts1
Completed observation slots8, spanning Day 0 and Day 1
Successful generation attempts / stored responses48 / 48
Coverage8 responses from each of the six providers
Future observation slots28 pending; none overdue at this check
Most recent stored responseSeptember 28, 2026, 6:25 AM Eastern

The runner’s every-minute schedule is active, and the three most recent inspected cron invocations succeeded. That confirms scheduler activity; the stored response records provide the stronger evidence of completed generation work. This report checks execution metadata, not the correctness of every answer.

Change since the previous report

This is the first published snapshot. There is no previous report for a measured comparison. Future reports will show changes in response coverage, checkpoint completion, evaluation coverage, review obligations, and findings supported by inspected evidence.

Review status

There are 36 evaluation records: 34 completed and two marked as requiring human review. No human-review records are stored. The most recent evaluation timestamp is September 27, so Day 1 evaluation coverage needs to be checked against the intended evaluation scope. “Completed” describes the recorded evaluator status; it is not a claim that Ben has independently validated the finding.

Limitations

One prompt cannot represent broad model capability. This audit has not adjudicated all raw outputs or validated the evaluator’s scoring. It does not establish provider rankings, correctness rates, or a longitudinal drift finding. Agreement and stability must be examined separately from correctness. The daily multi-model Council remains unverified and is outside this report’s observed experiment.

Leadership implications

The useful operating lesson is already visible: a scheduled task, stored outputs, automated assessments, and completed human review are separate states. A team needs to see each of them. The research architecture should preserve those distinctions so progress cannot quietly substitute for evidence.

As the study matures, the practical question will be whether the behavior a workflow depends on remains acceptable—and whether the people operating it can detect and respond to meaningful changes.

Next checkpoint and follow-up

Day 7 begins Sunday, October 4 at 5:54 AM Eastern, with repeats at 5:59, 6:09, and 6:24 AM. The next report is scheduled for November 1, followed by a new report every 30 days. Outstanding work includes the two flagged human reviews and verification of Day 1 evaluation coverage. If a future reporting window contains no new observations, its report will state that rather than manufacture a trend.