AI & Research

Can ChatGPT Predict Stock Prices From News? What the Research Actually Shows

A 2026 study finds GPT-4 reads the market impact of company news surprisingly well, but the realistic result is far weaker than the headline numbers, and it is not a ready-to-use trading strategy.

In short

  • GPT-4 was very accurate at reading whether company news was good or bad for a stock in the short term, but most of that accuracy describes a price move that has already happened before a normal trader could act.
  • The headline figures of about 93% (overnight) and 89% (intraday) are portfolio-day hit rates for the initial reaction, not a 90% individual-trade win rate.
  • The realistic result is the smaller continuation after the reaction: a 55 to 58% portfolio hit rate that faded after roughly one to two trading days.
  • The apparent edge collapses under trading costs (unprofitable at a 20 bps round-trip in the paper's illustration) and weakened sharply over the sample, from a Sharpe of 6.54 to 1.22.
  • The paper is evidence about how markets process information and underreact, not a ready-made trading strategy, and the authors say so explicitly.

Large language models are already used to summarise filings, classify sentiment and process huge amounts of text. A more ambitious question is whether they can spot information that has not yet been fully priced in. Alejandro Lopez-Lira and Yuehua Tang examine exactly that in their working paper Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models. Their results are striking: GPT-4 was highly accurate at identifying the direction in which markets initially reacted to company-specific news headlines.

But the most eye-catching result is also the easiest to misread. The paper does not show that a trader can feed a headline into ChatGPT, get a direction, and reproduce a 90% win rate. Most of that roughly 90% figure refers to a market reaction that has already occurred by the time a normal participant could realistically enter. The more relevant result is the smaller amount of continuation that occurs afterwards, which is statistically interesting but much weaker, highly sensitive to trading frictions, and explicitly not presented as an optimised strategy. That distinction is the whole point of this article.

1. The central distinction, in one picture

Almost everything that goes wrong in the retelling of this paper comes from collapsing two very different numbers into one. On one side is how well GPT-4 read the initial market reaction to news. On the other is how much of a stock's move was actually left to capture after that reaction. The first is huge and mostly unavailable; the second is modest and fragile.

Bar charts comparing the near-90% initial-reaction portfolio hit rate with the 55 to 58% post-announcement drift portfolio hit rate
Figure 1. Portfolio-day hit rates for the initial market reaction (left) and the post-announcement drift (right), for overnight and intraday news. Recreated from the numerical results reported in the paper.

2. What the researchers studied

Lopez-Lira and Tang investigate whether general-purpose language models can understand the short-term economic implications of firm-specific news. Their dataset contains roughly 159,000 firm-headline-date observations covering 4,123 US common stocks between October 2021 and May 2024, drawn from major US exchanges and matched with RavenPack data to keep the news relevant.

The period matters. The GPT-4 version used in the study had a stated knowledge cutoff in September 2021, so the researchers evaluate it on news released after that cutoff. That reduces the risk that the model is simply reproducing memorised historical outcomes. Each headline was shown to GPT-4 with a simple task: decide whether it was positive, negative or uncertain for the company's stock price in the short term. The answer was converted into a score of +1 for positive, 0 for neutral or uncertain, and −1 for negative, and then compared with actual price behaviour.

That setup allowed two very different questions to be studied separately: did GPT-4 correctly identify how the market reacted immediately, and after that reaction, did the stock keep moving in the same direction? Keeping those two apart is the whole point, because mixing them up is how a modest finding turns into a viral overstatement.

3. Finding 1: GPT-4 read the initial reaction very well

The headline result is that GPT-4's assessments lined up closely with the market's immediate response. For overnight news, the long-short portfolio built from GPT-4 scores produced a 93.3% daily portfolio hit rate for the initial reaction. For intraday news, the equivalent figure was 88.8%.

Two cautions come with those numbers. First, they are portfolio-day hit rates, not the share of individual headlines classified correctly. The metric asks whether the aggregate long-short portfolio moved in the predicted direction on a given day. Second, and more important, the initial move is largely not available to a normal trader once the news is out. For overnight news the initial reaction is measured from the previous day's close to the next day's open, so the stock has already repriced before the market opens.

News timingAvg. initial reaction, positiveAvg. initial reaction, negative
Overnight+1.27%−1.79%
Intraday+1.33%−3.11%

These results show that GPT-4 can interpret the economic meaning of company news. They do not show that those immediate returns were realistically available once the signal became observable. That is the first major thing to keep straight.

4. The realistic result: post-announcement drift

The more practically relevant part of the study is what happens after the first reaction. If markets were perfectly efficient, a public headline would be fully incorporated at once, and there would be no systematic continuation tied to the same information. Instead the researchers find a modest amount of continued movement in the direction GPT-4 predicted.

MeasureOvernight newsIntraday news
Portfolio hit rate58%55%
Mean return0.34%0.50%
Annualised Sharpe (pre-cost)2.972.63

A 58% or 55% portfolio hit rate is a very different claim from predicting stocks correctly nine times out of ten. The authors read the remaining drift as evidence of underreaction: GPT-4 appears to grasp some implications immediately, while the broader market prices the full information more gradually. The continuation is short-lived. For overnight news the average long-short return was about 34 basis points on the first trading day and 19 basis points on the next, after which the effect was no longer statistically meaningful. So the evidence points to a brief information-processing delay, not durable forecasting power over long horizons.

5. Negative news produced substantially stronger drift

One of the most interesting results is the asymmetry between good and bad news. For overnight news, the positive side of the portfolio produced roughly 8 basis points per day with an annualised Sharpe of 0.78, while the negative side produced about 26 basis points per day with a Sharpe of 2.01. In other words, much of the paper's return predictability came from stocks tied to negative news.

The authors connect this to the literature on limits to arbitrage. Negative information is harder to act on because betting against an overvalued security often requires short selling, which brings borrowing costs, availability constraints and greater implementation friction. As a result, negative information may take longer to be fully reflected in prices. Again, this is an explanation of the empirical result, not a suggestion to short stocks after bad news.

6. The effect was stronger in smaller stocks

The paper also finds that post-announcement predictability is stronger among smaller companies. GPT-4's ability to read the initial implication of news appears across company sizes, but smaller stocks show greater subsequent drift, which the authors again read as slower information incorporation.

The intuition is familiar. Large, liquid stocks are followed by more analysts, institutions, quant firms and market makers, so when material news appears there are more sophisticated participants racing to price it. Smaller companies get less attention and can face wider limits to arbitrage. The effect does not vanish entirely when very small and low-priced stocks are removed, but the strongest underreaction is concentrated in the less efficiently processed parts of the market. The useful distinction is that GPT-4 does not necessarily create the inefficiency; it may simply identify information that some parts of the market process more slowly.

7. Some news is priced fast, other news slowly

A neat contribution of the paper is grouping headlines by topic and comparing GPT-4's initial alignment with the later drift. For relatively transparent, quantifiable information the market reacts quickly and little drift follows. For information that needs more contextual reasoning, the stock tends to keep moving in the predicted direction afterwards.

Priced quickly (little drift)More delayed incorporation (more drift)
Earnings and revenue announcementsInsider stock transactions
Strategic partnership announcementsDividend announcements
Clinical-trial announcementsSpecialised healthcare conference information

The authors read this as evidence that market efficiency depends partly on the complexity of the information-processing task. A simple earnings surprise is rapidly turned into a numerical signal by existing participants and algorithms; a headline that needs more contextual reasoning takes longer to digest.

8. More capable language models performed better

The paper compares GPT-4 with a range of other models, including GPT-3.5, BERT-based systems, Llama models and FinBERT. A clear ranking appears for the overnight drift strategy, measured by annualised Sharpe before transaction costs.

Horizontal bar chart of annualised Sharpe ratios for GPT-4, GPT-3.5, DistilBART-MNLI, BART-Large and Llama2-70B
Figure 2. Reported pre-cost Sharpe ratios of the overnight drift strategy across models. GPT-4 leads by a wide margin. Recreated from the paper's reported values.
ModelDrift-strategy Sharpe (pre-cost)
GPT-42.97
GPT-3.51.66
DistilBART-MNLI1.26
BART-Large1.05
Llama2-70B0.97

Simpler models such as GPT-1, GPT-2 and BERT generally did much worse and showed no reliable positive drift prediction. FinBERT is a telling case: it was designed for financial text and did well at reading the initial reaction, but its drift performance was much weaker. The authors argue GPT-4's advantage comes from broader semantic and contextual reasoning rather than simple sentiment detection. The relevant question is not just "is this headline positive or negative?" but closer to "what does this specific event mean economically for this specific company, and how is the market likely to interpret it?" That is a harder problem than sentiment detection, which is why the more capable models pulled ahead.

9. The backtest measured predictability, not a tradable strategy

The paper also builds a simple daily long-short portfolio from GPT-4's classifications. Before transaction costs, the cumulative result is enormous, roughly 700% between October 2021 and May 2024. It would be easy to read that as an extraordinary trading strategy, but that is not what it is meant to be. The authors state plainly that the portfolios are not optimised for implementation; they exist to measure GPT-4's raw forecasting power. The ~700% is a way of expressing how strong the predictive signal was, not a claim about money a trader could have taken home.

That is worth sitting with, because the portfolio is built in a way no one would actually trade. It needs very high turnover, about 190% daily turnover in the baseline equal-weighted version, so if anyone did try to run it as a live strategy, trading costs would erode the result quickly.

Bar chart showing cumulative backtest return falling from about 700% at 0 bps to unprofitable at 20 bps
Figure 3. Approximate cumulative backtest return under different assumed round-trip costs. At about 20 basis points the illustrative portfolio stops being profitable. These are backtest results from the authors' illustrative portfolio, not expected future returns.
Assumed round-trip costApproximate cumulative result
0 bps~700%
5 bps>300%
10 bps>100%
20 bpsUnprofitable

And real execution involves more than commissions: price impact, short-borrow costs, slippage, liquidity limits and the difficulty of filling many positions at assumed prices. None of this makes the paper's number wrong or misleading, because the number was only ever a measure of predictability. It simply marks the gap between return predictability, which the paper demonstrates, and an implementable, scalable strategy, which the authors never claim to have built and do not try to.

10. The apparent edge weakened sharply over time

Perhaps the most important result for anyone tempted to treat this as a current strategy is how much the performance decayed. The reported annualised Sharpe ratio of the overnight GPT-4 strategy fell steadily across the sample.

Line chart showing annualised Sharpe ratio falling from 6.54 in 2021 Q4 to 1.22 by Jan-May 2024
Figure 4. Reported pre-cost Sharpe of the overnight strategy over the sample. The authors read the decline as suggestive evidence of increasing market efficiency, though they do not claim it as proven causation. Recreated from the paper's reported values.
PeriodAnnualised Sharpe ratio
2021 Q46.54
20223.68
20232.33
Jan–May 20241.22

The authors suggest that wider adoption of language models may be contributing to greater market efficiency. The logic is simple: if only a few participants can rapidly interpret complex information, they hold an informational edge, but if millions gain similar capabilities, that information is priced in faster. The very technology that spots the inefficiency may help erase it. The authors are careful about causality, though: changing market conditions or other factors could also explain part of the decline, so the evidence is consistent with LLM adoption improving efficiency without proving it caused the drop. Either way, a decaying edge is another reason not to treat the study as a ready-made system.

11. What the paper does not show

The research is compelling, but several popular claims go well beyond the evidence.

The claimWhy the paper does not support it
ChatGPT predicts individual stocks with 90% accuracyThe ~90% figure is the direction of an aggregated portfolio's initial reaction, not a 90% individual-trade win rate.
The 90% reaction was tradableFor overnight news the reaction is measured from prior close to next open, so most of it is already priced before a normal entry.
The paper gives a complete strategyIt does not optimise execution, risk, sizing, leverage, liquidity, short availability or market impact, by the authors' own statement.
The historical edge will persistThe reported Sharpe fell substantially over the sample; anomalies can shrink as participants find and exploit them.
Statistical predictability equals realisable returnsHigh turnover and cost sensitivity mean the two are not the same thing.
Do not confuse forecasting evidence with a trading strategy. The paper documents statistically significant return predictability, but the authors do not provide an optimised implementation. The results depend on portfolio construction, transaction costs, short-selling assumptions, liquidity and execution, and the strongest headline accuracy relates to the initial price response, much of which occurs before a normal post-news trade could be placed.

12. Why the paper still matters

Once the exaggerated reading is stripped away, the paper arguably gets more interesting. Its strongest contribution is not that an AI chatbot can replace a trader. It is evidence about how markets process information: advanced language models can extract economically meaningful signals from short headlines; markets price some information almost instantly; more complex information is incorporated gradually; delayed incorporation is stronger where arbitrage is harder; semantic reasoning seems to matter beyond simple sentiment; and better information-processing technology may itself raise market efficiency.

That places the study in a broad literature on underreaction, information diffusion and limits to arbitrage. The language model is useful not only as a forecasting tool but as a research instrument: by comparing what GPT-4 understands immediately with what the market prices immediately, researchers can flag the kinds of information investors find hardest to process in real time.

The safest one-line summary is this: GPT-4 was highly effective at identifying the economic direction of company news, but most of the roughly 90% initial-reaction result was already embedded in prices before a normal trader could act, and the remaining predictability was much smaller, around a 55 to 58% portfolio hit rate, and sensitive to costs, size, news type and time period. That is an important result about markets. It simply is not a ready-to-use trading strategy, and whether any statistical edge can be turned into an implementable one is a separate question this paper does not try to solve.

Test a strategy against the odds, not the hype

Headline accuracy figures rarely survive contact with costs, sequencing and hard limits. If you want to see how a set of assumptions actually plays out over thousands of simulated paths, run it through the free pass rate simulator.

Open the pass rate simulator →
or

Read more research explainers written for traders, not for headlines.

Browse all research explainers →

Frequently asked questions

Can ChatGPT predict stock prices from news headlines?

The research shows GPT-4 is very good at classifying whether a headline is good or bad news for a company in the short term. That is not the same as predicting a tradable price move. Most of the eye-catching accuracy describes a reaction that has already happened in the price before a normal trader could act, and the smaller effect that remains afterwards is weaker and sensitive to trading costs.

What does the 90% hit rate actually mean?

It is a portfolio-day hit rate for the initial market reaction, not a 90% win rate on individual trades. For overnight news the reaction is measured from the previous close to the next open, so most of that move is already in the price before the market opens.

What is post-announcement drift?

It is the tendency of a stock to keep moving in the same direction for a short period after news is released, rather than repricing fully and instantly. In the study this continuation was much smaller than the initial reaction, roughly a 55 to 58% portfolio hit rate, and it faded after about one to two trading days.

Why did the negative-news side of the portfolio perform better?

The authors link it to limits to arbitrage. Acting on negative information often requires short selling, which involves borrowing costs, availability limits and other frictions, so negative information can take longer to be fully reflected in prices. This is an explanation, not a suggestion to short anything.

Does the paper give a trading strategy?

No. The authors state that their portfolios are built to measure GPT-4's raw forecasting power, not to be an optimised, implementable strategy. They do not model execution, sizing, leverage, liquidity, short availability or realistic market impact.

Why do the returns collapse once transaction costs are added?

The strategy needs very high turnover, with about 190% daily turnover in the baseline portfolio. At that turnover, small per-trade costs compound quickly. In the paper's illustration a 20 basis point round-trip cost is enough to make the approach unprofitable.

Did the edge get weaker over time?

Yes. The reported pre-cost annualised Sharpe ratio of the overnight strategy fell from about 6.54 in late 2021 to about 1.22 by early 2024. The authors suggest wider access to language models may be making markets more efficient, though they are careful not to claim that as the only cause.

Does this paper apply to prop-firm challenges?

Not directly. It studies a large cross-section of US stocks over a specific period. It is useful as background on how markets process information and why headline accuracy figures can be misleading, but it does not describe any prop-firm rule set or evaluation.

References

Lopez-Lira, A., & Tang, Y. (2026). Can ChatGPT forecast stock price movements? Return predictability and large language models [Working paper]. SSRN.

Figures were recreated from the numerical results reported in the working paper and are not reproductions of the paper's own charts. Working-paper results can change between versions; verify against the latest manuscript.

Educational and informational only. This article reviews academic research and does not recommend any security, strategy, position size, product or prop firm. Historical strategy statistics shown here are the authors' reported backtest results before or after their stated transaction-cost assumptions, not expected future returns.
← All blogs

Disclaimer: DanFin is provided for educational and informational purposes only and does not constitute financial, investment, or trading advice, nor a recommendation of any firm, product, or strategy. The simulator and calculators are simplified statistical models based on the figures you enter; their outputs are hypothetical, are not predictions, and do not guarantee future results. Trading leveraged products carries a substantial risk of loss and is not suitable for everyone. Do your own research and consider consulting a licensed professional before making any financial decision. Some links on this site are affiliate links; if you use them, DanFin may earn a commission at no extra cost to you, which never changes the results the tools give you or the content shown.