Prompt engineering in Amazon Quick determines how accurately and reliably the platform’s AI-powered features respond to your natural-language requests. Whether you’re building custom agents, authoring automation flows, or querying data through conversational analytics, the way you structure your prompts directly shapes the quality of the output you receive. In this post, you will learn the

AI forecasters are catching up to humans. It’s a new opportunity for businesses—if they hire the humans too. | IBM
Yann Rivière is as close as someone can get to a real-life fortune teller. He’s accurately forecasted rocket launches, the outbreak of global conflicts and the price of cryptocurrency, all months in advance, even when the outcomes had seemed unlikely to his peers.
Rivière is among a small slice of people who can predict the outcome of future events with unusual accuracy and precision. He competes on Metaculus, a forecasting platform that hosts prediction tournaments where competitors try to predict everything from chess championships to gas prices to the whereabouts of individual sharks. Some have prize pools as high as USD 50,000, while others offer no incentive aside from bragging rights and the love of the game. Out of tens of thousands of players, Rivière consistently ranks near the top, with an all-time tournament performance in the 99.99th percentile.
But lately Rivière has some stiff competition, as a new crop of competitors rises through the ranks. Backed by independent tinkerers and well-funded startups alike, AI models like Preseen, Cassi and Laertes are catching up to Metaculus’ top forecasters. Bots have already bested individual participants, including Rivière in some tournaments. In one recent competition, where participants forecast across dozens of wide-ranging questions, AI models earned the first- and second-highest scores—a tournament first.
“The first time, it hurt. It’s like, ‘Okay, I’ve been beaten by a robot.’” Rivière told IBM Think in an interview. “Especially since it’s not even like I had a bad tournament or something. They were clearly in the top 1%. They’re just that good.”
Metaculus’ tournament results reflect a larger pattern. Multiple benchmarks, from ByteDance’s FutureX to the University of Chicago’s Prophet Arena, show that AI’s forecasting capabilities are quickly improving. The Forecasting Research Institute (FRI), meanwhile, declared in July that its leaderboard’s highest-ranking models are now “statistically indistinguishable” from elite human forecasters. This rapid progress has set off a race among AI startups to design the first model that can consistently beat elite-level predictors across any subject—what might be considered forecasting’s equivalent of artificial general intelligence (AGI).
But the industry’s scramble to build an AI oracle renews a debate about the value of forecasting inside businesses. Rivière, who works at AI startup Mantic as a professional forecaster, is the rare outlier who’s turned his talent into his job. Elite forecasters might wield their powers during tournaments or try their hand at prediction markets like Kalshi and Polymarket, but relatively few are helping executives make decisions. One reason is that it can take hours for a forecaster to make even a single forecast, and by then, that forecast might already be outdated. As a result, businesses would need to hire an army of elite predictors to keep up with demand, which often isn’t feasible at scale.
For quantitative forecasts, such as financial planning, demand forecasting and predictive maintenance, enterprises have historically turned to time series models, machine learning models that work with numbers instead of words. For open-ended, strategic questions, such as whether a competitor will launch a new product next year, enterprises have traditionally relied on executives, who filter expert opinions and customer feedback through their own experience and intuition.
Now, as AI approaches parity with humans, that balance is starting to shift. But forecasters warn that the answer isn’t as easy as throwing any question at a frontier model and expecting an accurate answer. The companies best positioned to benefit are pairing AI’s breadth and speed with competitive forecasting’s processes, theories and insights—even as their competitors try to get blood from a stone by running AI with no human in the loop.
Ironically, the rise of AI forecasters might make human predictors more attractive hires. To come up with an accurate prediction, forecasters must first build a theory for why one outcome is more likely than another, separating key variables from irrelevant noise. That process, which businesses are increasingly looking to replicate with AI, can yield insights that prove more consequential than the prediction itself. Human forecasters can provide a blueprint for how teams might use AI to gather evidence, experiment with unconventional ideas and challenge long-held assumptions.
Inside the mind of an elite forecaster
Like many of his peers, Rivière fell into forecasting by accident. Metaculus’ prediction platform was one of the “many weird places on the internet” he stumbled upon while creating longform video essays on YouTube, where he’s covered topics such as film studies, geopolitics and the search for extraterrestrial life. Around 2021, while studying for a law degree, he created a Metaculus account and made a few predictions on a whim. He gradually climbed the leaderboard, and by 2026, he’d joined Metaculus’ roster of pro forecasters and landed forecasting gigs with RAND, an influential think tank founded in 1948, and the Swift Centre, a forecasting-focused advisory and research organization.
Rivière pins his success less on encyclopedic knowledge and more on his ability to check his own biases, paired with his insatiable curiosity. “I’m way too curious for my own good,” he said.
For example, to predict whether Shakira’s 2026 World Cup official song would out-peak “Waka Waka (This Time for Africa),” the hit she wrote for the 2010 World Cup, he weighed numerous factors, including: that the Billboard Hot 100 chart operates on a slight delay, that the US co-hosted the World Cup this year, that the opening and closing ceremonies might impact streams and that this year’s song had a higher potential for virality because of the proliferation of the internet since 2010—all of this analysis for just one of the thousands of predictions he’s logged on the platform.
This attention to detail is common among the world’s top forecasters, according to University of Pennsylvania psychologist Philip Tetlock, who trademarked the term “superforecaster” to refer to this distinctive group. Tetlock’s 2015 book Superforecasting: The Art and Science of Prediction, which he cowrote with author Dan Gardner, argues that when it comes to forecasting accuracy, IQ and subject matter expertise aren’t enough. In fact, knowing too much about a topic can magnify one’s blind spots.
One oft-cited experiment found that a group of superforecasters performed with a roughly 30% lower error rate than intelligence officers across a range of geopolitical predictions—even though the officers had exclusive access to classified information. (Other researchers have contested the claim that superforecasters consistently outperform domain experts.)
Accurate forecasting requires a blend of analytical and abstract thinking, along with open-mindedness and a willingness to admit uncertainty. Expert forecasters are especially skilled at recalibrating their beliefs as new evidence emerges. And crucially, they can admit when they don’t know something and can pinpoint the degree to which they’re unsure.
To formulate a new forecast, Rivière starts by establishing a base rate, (a foundational probability rooted in historical data). If he wants to predict, say, whether a first-time champion will win the next World Cup, he’ll first determine how frequently that scenario has unfolded in the past.
Then, he’ll look at the particulars of that year’s competition—by analyzing the top contenders, evaluating the health of each player and so on—to further home in on the likeliest winners. He tries to spot salient themes and patterns across articles, interviews and studies, while at the same time filtering out irrelevant or unreliable data. He sometimes does his own reporting, posing questions on niche forums and, in the case of a forecast on competitive Rubik’s Cube solving, chatting with enthusiasts to get an on-the-ground perspective.
The last step, though, can be difficult to put into words. “At the end, once you’ve got all of that [research], there’s some vibe aspect to it,” Rivière said. “You have to go with your gut.”
AI has historically struggled to make that final intuitive leap. But in the past year, against the odds, that’s started to change.
How to beat a frontier model at forecasting
Alongside his typical forecasting work, Rivière has lately become a mentor. But instead of teaching humans, he’s guiding LLMs toward more accurate responses for AI startup Mantic. Rivière reviews model responses and tries to identify their missteps, such as when they overlook a crucial piece of information during retrieval. The goal: designing an architecture that transforms LLMs into forecasting experts.
Mantic is one of perhaps dozens of companies that have attracted VC interest and enterprise customers in recent months with the promise of forecasting future outcomes at or above a human level, but faster and cheaper than would be possible with humans alone.
To get a sense of the industry’s growth, consider ForecastBench, the prediction accuracy benchmark hosted by Tetlock’s Forecasting Research Institute (FRI). Of the 35-plus organizations represented on the leaderboard, more than half are smaller, prediction-focused startups and independent developers competing with the likes of Google, Anthropic and OpenAI.
Researchers have yet to coalesce around a single strategy for designing a winning model. But one way to fast-track performance is to start with a powerful base LLM, according to Nick Merrill, a Senior AI Research Consultant at FRI. “As models become more generally sophisticated, they get better at superforecasting—and they get better at everything else,” Merrill told IBM Think. That trend suggests that reasoning and prediction prowess are closely related, even if the precise mechanisms that drive forecasting accuracy remain elusive to developers.
But a recent Metaculus report found that smaller companies and even individual developers can narrow the performance gap by applying custom prompting and training techniques to slightly smaller or less powerful models. Effective strategies include running forecasting simulations with historical data, performing agentic searches across news sites and compiling multiple model responses into a consensus forecast. Combined, these approaches can be “worth about nine months of base model progress,” according to the analysis.
Judging by their recent tournament scores, forecasting models seem to excel at answering the kinds of clear-cut, well-defined questions that researchers track on competitive forecasting platforms; it’s harder to measure how they might perform in real-world business scenarios (although they’re already showing some success in specific categories, such as predicting which technology ventures will succeed, with higher accuracy than human experts).
Still, businesses often want to know how best to prepare for the future, and better yet, how to mold the future in their favor, which a probability can’t provide. That helps explain why even the CEO behind one of Metaculus and ForecastBench’s highest-ranking AI models is unconvinced that accuracy alone will give his clients an edge.
A goal more important than accuracy
As an intelligence officer in the Royal Air Force and an Expert Advisor to the Prime Minister in the UK, Keith Dear learned that an accurate forecast is only as valuable as an organization’s ability to use it. Now, as the CEO of forecasting startup Cassi AI, Dear challenges enterprises to focus less on arbitrary probabilities and more on pinpointing what they want to achieve.
“For most customers, a 3% better forecast is not the thing holding them back,” Dear told IBM Think. “Being able to spot the things that are most likely to change a future outcome and then [concentrating] their resources on that will be much more valuable.”
Another issue is that aside from high-stakes decisions—when to expand into a new market, say, or whether to adopt a remote-first work policy—organizations don’t always have the bandwidth to apply rigorous research to every problem. For example, a 2025 Salesforce report found that fewer than half of business leaders feel confident they can “use data to drive action and decision-making” in their organization. Too often, the alternative is a reliance on human-constructed narratives rather than insights drawn from firm evidence, according to Dear.
“Prediction at the moment is all implicit and story-based,” he said. “Decisions are made around the table by arguing your case through social influence.”
AI has upended that dynamic. Just as an elite human forecaster aims to adopt a clear-eyed view of the data, platforms like Cassi can spot hidden variables that guide teams toward a more evidence-based perspective—and continuously recalibrate those estimates as new information emerges.
For example, when the UK Parliament recently asked Dear about the likelihood that the nation would face a major conflict by 2036, he posed the question to Cassi. Aside from delivering a probability (20%), Dear reported to Parliament, Cassi highlighted some of the biggest factors that might shift the outcome, such as defense spending, alliances and AI advances—all factors that Parliament can influence.
Findings like those can help remind organizations that they aren’t passive observers; their actions often play a role in determining what happens next. Cassi hints at a future where explainability, or helping users understand the drivers and assumptions behind an AI-generated forecast, might come to matter as much as sheer accuracy.
Developers are also starting to infuse quantitative models—which are traditionally highly accurate but hard to interpret—with LLMs so that users no longer have to negotiate between precision on the one hand and explainability on the other.
For example, IBM’s Planning Analytics platform combines statistical forecasting with a generative AI layer to make forecast results easier to act on. “With machine learning-powered statistical models, we generate the numbers,” Planning Analytics Product Manager Svetlana Pestsova told IBM Think. “With LLMs, we interpret the results in human language.” That means users don’t waste time interpreting the results and can instead focus on how to respond. “Instead of spending hours or days or months producing the forecast, you can concentrate on decision-making,” Pestsova said.
But even as prediction platforms make generating forecasts easier, embedding them inside an enterprise is an entirely different challenge. Teams must decide which AI suggestions to apply and which to ignore, which metrics to measure progress with and where to set guardrails—subjective problems that humans can’t delegate to AI.
Self-fulfilling prophecies
While frontier models can expand an organization’s forecasting capabilities, one overlooked challenge is integrating those tools into existing business workflows and management, governance and auditability frameworks, Romain Le Duc, Principal Data Scientist at IBM Consulting, told IBM Think in an interview.
Le Duc helps teams build end-to-end AI forecasting systems with a focus on integration and orchestration, where multiple models and agents work together and draw on proprietary data alongside external sources. “A standard ChatGPT session will not do that out of the box because it does not integrate the context of your data, your company’s operating processes and policies and all those customizations that make you better than your competitors,” Le Duc said.
That additional context, he said, also helps make a forecast explainable, and as a result, contestable and reversible. Transparent systems enable planners to see which factors drove a forecast and allow auditors to trace how it was produced. “Explainability, transparency, security and fairness cannot be bolted on at the end,” Le Duc said. Instead, the system must be designed from the ground up with those principles in mind.
At a more fundamental level, though, many organizations are struggling with the ethical implications of handing more responsibilities to AI. According to the IBM Institute for Business Value, over half of executives cite ethics as a major barrier to AI deployments. That starts with how LLMs can reshape our sense of inevitability. One recent paper by European researchers found that people who receive incorrect information from chatbots become more confident that they’re right and less likely to consider alternative possibilities.
But uncertainty is valuable, helping forecasters take a more cautious or deliberate approach when they might otherwise act too aggressively, according to Rivière, the Mantic forecaster. “Even when we understand the subject pretty well, we’re able to admit that sometimes we’re just wrong,” Rivière said. “I would say the most important thing is to be able to change your mind.”
The act of forecasting itself can shift the outcome of an event, writes Carissa Véliz, an Associate Professor at the Institute for Ethics in AI at the University of Oxford. “Our expectations bend the social world toward our predictions,” she wrote in her 2026 book, Prophecy: Prediction, Power, and the Fight for the Future, from Ancient Oracles to AI. “When someone forecasts that the world will be a certain way, they are commanding that others obey their wishes and bring that world about.” For example, when an epidemiologist alerts the public of an impending pandemic, this warning might convince governments and communities to mobilize against it, reducing its severity. When an economist suggests that a country might soon face a financial crash, investors might be incentivized to withdraw their stakes. And when nations fear war, they’re more likely to prepare for it, which can further escalate tensions.
Organizations, and society at large, are still figuring out where to draw the line. Most people don’t want an AI model to arbitrarily decide whether they qualify for health insurance or a mortgage or a job. In fields such as healthcare and insurance, regulators might consider requiring companies to use forecasting platforms strictly for administrative tasks—or help define what counts as human oversight to prevent the “human in the loop” from becoming a rubber stamp for AI. For decisions that affect people, transparent and contestable guidelines can help ensure that everyone is playing by the same rules and that stakeholders can openly challenge criteria they find unfair or objectionable, Véliz argues.
There are other potential risks: without robust guardrails, AI can serve as a convenient scapegoat when a company or individual makes a bad decision, reducing accountability. And in applications requiring innovation, the algorithms’ tendency to suggest the safest option can stifle ingenuity and disincentivize experimentation.
While the days of human-dominated forecasting tournaments might be coming to an end, human judgment and critical thinking will remain essential in enterprise settings, even as our relationship with AI continues to evolve. Lately, Rivière has stopped thinking of LLMs as supercharged search engines and more as competent colleagues who can engage in thoughtful debates with him. Older LLMs “helped your research, but they didn’t really think, or at least think well, about what everything meant,” Rivière said. Now, AI can synthesize disparate pieces of data, weigh what matters most and evaluate the consequences of different actions in addition to generating a probability.
Where humans have the undisputed edge, Rivière said, is in pinpointing what to forecast, what parameters to analyze further and, perhaps most importantly, why a particular forecast matters in the first place.
The latest AI trends, brought to you by experts
Get curated insights on the most important—and intriguing—AI news. Subscribe to the twice-weekly Think Newsletter. See the IBM Privacy Statement.
