Writing / AI and Automation

What Happens When You Ask Six AI Models the Same Question for Six Months?

Ben Tartaglia · Published October 2, 2026

Ask six AI models the same question and you may get six polished answers. The useful question is what survives inspection—and whether it continues to survive when you ask again.

That’s the question behind the longitudinal experiment in The Glowing Oracle Lab. The six-month horizon is part of the design. We’re at the beginning of that timeline, not presenting a completed six-month study.

The first live cohort began on September 27, 2026. Its prompt describes a 120-meter steel pedestrian bridge at 10°C, gives a linear expansion coefficient of 12 × 10⁻⁶ per °C, and asks how much the bridge expands when the temperature reaches 35°C. It requires the formula, calculation, answer in millimeters, and assumptions, without web browsing.

Under those supplied conditions, the reference calculation is straightforward: 120 × 12 × 10⁻⁶ × 25 = 0.036 meters, or 36 millimeters. That makes the task useful as a controlled starting point. We can distinguish arithmetic and instruction following from persuasive presentation.

The participating providers are OpenAI, Anthropic, Google, xAI, Mistral, and DeepSeek. Each checkpoint includes fresh sessions at minute 0, 5, 15, and 30. The stored schedule spans Day 0, 1, 7, 14, 21, 30, 90, 180, and 365.

The short repeats help establish ordinary variation around a checkpoint. The later repeats create opportunities to observe changes over time. Those are different comparisons. A different answer fifteen minutes later doesn’t, by itself, demonstrate a provider update or a lasting behavioral shift.

As checked on October 2, Day 0 and Day 1 had produced 48 stored responses: eight per provider. There were 36 evaluation records, including two requiring human review. No human reviews were recorded. These counts describe execution and review status; they don’t establish a winner or a finding about drift.

The Lab’s rubric includes factual accuracy, instruction adherence, assumption control, reasoning quality, uncertainty calibration, and content drift. I want those dimensions kept visible. A response can change substantially and become more useful. It can also remain remarkably consistent while repeating an error. Stability and correctness are separate questions.

Original prompts and outputs matter because a summary can erase the very detail we need to inspect. Model identifiers, settings, timestamps, tool policies, and evaluator versions matter because a comparison needs a record of what changed around the answer.

There are limits. One quantitative prompt cannot represent every use case. An automated evaluator needs scrutiny of its own. Agreement between providers isn’t independent proof of correctness. Repeated samples can reveal patterns, but attributing a pattern to a specific cause requires more evidence.

The leadership application is practical: test the behavior your workflow depends on, preserve enough evidence to investigate it, and repeat the test after meaningful changes.

Over the coming checkpoints, I want to learn whether the material answer, stated assumptions, and adherence to instructions remain dependable. The value of this experiment will come from careful observation—including observations that are less exciting than the headline.