Artificial Intelligence has transformed how we search for information, generate text, and analyze data. But can it accurately predict the future? ABC News In-depth recently conducted a fascinating experiment to find out, exploiting a built-in limitation of modern AI to test its foresight.
The results reveal a technology that is incredibly adept at understanding complex geopolitical strategies, yet surprisingly terrible at guessing the stock market or who will win a sports championship.
The “Frozen in Time” Flaw
To understand how AI predicts the future, you first have to understand how it learns. Training an AI model involves feeding it massive datasets scraped from the internet. Once this training phase ends, the AI’s knowledge base is essentially locked.
Think of it like the protagonist in the movie Memento, who can no longer form new long-term memories after a head injury. An AI model is frozen on the day its training finishes. If you ask it about an event that occurred after that cutoff date—such as a sudden geopolitical crisis or the assassination of a world leader—it will draw a blank.
While this lack of current awareness can be frustrating for a user trying to catch up on the news, it presents a unique scientific opportunity. What if you talk to an AI “in the past” and ask it to predict an event that has already happened in your present?
The Experiment: Operation Epic Fury
To test this, the ABC News team set up an experiment using three prominent AI models: Claude, Grok, and ChatGPT. Because the models had training cut-offs that predated a major recent geopolitical event—a hypothetical military strike on Iran dubbed “Operation Epic Fury”—they had no knowledge of how the situation actually unfolded.
The AI models were given a simulated intelligence dossier and battle plan matching what the administration had on its desk before the strike. Their web search capabilities were disabled, ensuring they couldn’t cheat by looking up the real-world outcome.
They were then asked to evaluate the likelihood of several stated objectives being achieved:
- Would the regime be toppled?
- Would the Iranian missile industry be destroyed?
- Would the Navy be annihilated?
- Would the Revolutionary Guard surrender?
Where AI Excels: Military and Geopolitical Strategy
The models performed with spooky accuracy. Against the actual historical outcomes, the AI consensus correctly predicted that only one or two of the primary objectives would be achieved.
When pressed on longer-term ramifications—such as whether a durable ceasefire would be reached, if oil prices would spike, or if the Strait of Hormuz would be closed long-term—the AI models correctly predicted 12 out of 14 outcomes.
Why is AI so good at this? AI models have consumed virtually every unclassified report published by war colleges, military strategists, and geopolitical think tanks. They essentially represent the ultimate consensus of human expertise. When dealing with rigid military doctrine and geopolitical cause-and-effect, AI acts as an incredible strategic analyst.
Where AI Fails: Sports, Stocks, and Human Chance
Before you try using an AI as a sports almanac to get rich, there’s a catch. The same experiment tested AI’s ability to predict outcomes driven by human chance, emotion, and fluid dynamics.
The models were asked to predict:
- Sports: The top finalists for the NRL, Premier League, NFL, and NBA.
- Finance: The price of Gold, Bitcoin, Oil, and the performance of major stock indexes like the S&P 500.
- Politics: The outcome of chaotic political leadership spills (like the recent replacement of the UK Prime Minister).
In these arenas, AI failed spectacularly. None of the models achieved better than 50% accuracy in sports predictions. They incorrectly predicted that the S&P 500 would drop when it actually rose, and completely misread Bitcoin’s trajectory. In politics, they struggled to predict unpredictable human choices, completely missing the unexpected rise of a dark-horse candidate for UK Prime Minister.
AI models suffer heavily from recency bias. Because they operate on patterns, they assume that whatever happened recently will continue to happen. If a sports team usually makes the finals, the AI assumes they will make it again. They cannot account for sudden injuries, market panic, or human irrationality.
The Irony of the AI Boom
We are currently living in an era where tech giants are pouring trillions of dollars into AI infrastructure. Data centers are consuming larger proportions of national budgets than the Apollo program or the Manhattan Project did.
Yet, there is a stark irony in this massive investment in artificial “experts.” When human leaders are presented with complex problems, they frequently ignore the consensus of both human analysts and machine intelligence.
As the experiment highlighted, when a leader is given data that contradicts their personal beliefs or desired outcomes, they often roll the dice and go with their “gut instinct.” AI doesn’t have a gut feeling, a political ideology, or an axe to grind. It simply offers the most mathematically probable outcome based on collective human knowledge.
AI might not be able to tell you which stock to buy or who will win the championship, but it is an incredibly powerful tool for analyzing complex, logical systems. The real question isn’t whether AI can predict the future, but whether we are actually willing to listen to it when it does.
Large Language Models (LLMs) learn to generate text by recognizing patterns across massive datasets. At their core is the Transformer architecture, which processes language mathematically rather than reading it linearly.
How Self-Attention Connects Data
When you input a sentence, an LLM breaks the text into pieces called “tokens” (words or sub-words) and converts them into numerical vectors called embeddings. To understand the context and meaning of those tokens, the model uses the Self-Attention Mechanism.
Self-attention allows the model to evaluate the importance of every token in a sequence relative to every other token, regardless of how far apart they are. For example, in the sentence “The bank of the river,” the word “bank” means something completely different than in “The bank on the corner.”
By calculating the mathematical relationship between “bank” and “river,” the model assigns a higher “attention weight” to that connection. This allows the AI to map out dependencies, understand pronoun references, and grasp the semantic meaning of a full paragraph simultaneously.
Here is a visual breakdown of how that architecture actually processes language:
Generating video…This may take a few mins
The Three-Stage Training Pipeline
Understanding the connections between words is only the baseline capability. Creating a usable AI assistant requires a specific, three-stage training pipeline where each step builds upon the last.
| Stage | What It Does | How It Learns | The Result |
| 1. Pretraining | Learns language, grammar, reasoning patterns, and world knowledge. | Self-supervised learning on internet-scale text. The model is simply trying to predict the next token in a sequence. | A “Base Model” that is highly knowledgeable but unpredictable. It can generate text but doesn’t know how to follow instructions. |
| 2. Supervised Fine-Tuning (SFT) | Learns how to behave like a helpful assistant and follow directions. | Trains on thousands of highly curated, human-written prompt-and-response pairs. | An instruction-following model that understands how to format answers and respond to queries. |
| 3. Reinforcement Learning from Human Feedback (RLHF) | Learns what humans actually prefer regarding safety, helpfulness, and tone. | Human testers rank multiple model responses from best to worst. A separate “reward model” learns these preferences to automatically score and optimize the LLM. | A polished, aligned model that is safe to deploy to the public. |


Leave a Reply