Large language models

  • 详情 Option Return Predictability via Large Language Models
    We investigate the capabilities of Large Language Models (LLMs) in generating novel alpha factors for option returns. Utilizing a structured prompt-engineering approach, LLMs like GPT-5 can directly create factors for two distinct options markets: the mature U.S. market and the emerging Chinese market. Empirical analysis further reveals that the LLM-generated factors exhibit remarkable and robust performance, delivering statistically signifcant returns in both all-sample and extensive out-of-sample tests. Beyond their statistical signifcance, such factors are economically meaningful. They display low self-correlation, indicating genuine innovation, and are grounded in sound economic rationale derived from market microstructure and behavioral fnance principles, showcasing a key advantage over traditional machine learning models.
  • 详情 Validated Corporate Narratives and Bank-Affiliated Investment: A Large-Language-Model Approach
    Technology firms are often financed on narratives about products, contracts, customers, and technological progress well before these developments appear in accounting statements. We ask when such narratives become economically informative. Our central idea is that narratives should matter more once they can be linked to later verifiable outcomes rather than treated as stand-alone text.Using listed Chinese technology firms, we develop a validated corporate narrative framework for bank-affiliated investment, a setting in which investors must screen with soft information ex ante and then monitor hard realization and downside risk ex post. We use GPT-5.1 to extract business claims from management discussion, investor-relations records, exchange Q&A, and earnings-roadshow materials, and to label later claim–evidence pairs as support, partial support, conflict, duplicate, or irrelevant. We then connect these labels to official announcements, procurement awards, permits, project updates, and negative-event disclosures to construct a validated firm-month signal. The broad merged panel contains 592 firms and 30,169 firm-month observations; the main return tests use 576 firms and 18,230 firm-month observations over 2022–2024. A simple production rule that combines a low-narrative-premium component with hard-narrative and hard-event anchors, together with a separate downside-risk gate, delivers an implementable annualized long-short return of 8.93% in bank-invested firms after trading costs. The signal is much weaker in non-bank firms, predicts future gross-margin improvement more strongly than future ROE, and improves downside screening.
  • 详情 When LLMs Go Abroad: Foreign Bias in AI Financial Predictions
    We document “foreign bias” in AI financial predictions, reversing the classic home bias. U.S.-based ChatGPT is systematically more optimistic than China-based DeepSeek about Chinese firms—in price predictions and directional forecasts—yet significantly less accurate. Evidence supports an information-availability mechanism: bias is strongest when U.S. media coverage of Chinese firms is limited and attenuates for cross-listed firms. Crucially, injecting Chinese news eliminates the prediction gap. Both models produce similar forecasts for U.S. firms, consistent with broader worldwide coverage. LLMs trained in different information environments can create divergent signals, with implications for investors and policymakers as AI increasingly intermediates global markets.
  • 详情 AI Narrative Gap as a Firm Characteristic: Analyst Over-Optimism and Return Reversals
    We propose the AI Narrative Gap as a novel firm characteristic—the systematic divergence between a firm’s AI strategic narrative intensity and its subsequent AI capital expenditure commitment—and document its capital market consequences. Using Chinese A-share listed firms from 2015 to 2022, we show that firms with a wider AI Narrative Gap attract significantly more optimistic and less accurate analyst earnings forecasts. These distorted expectations, in turn, predict lower subsequent stock returns, lower industry-adjusted abnormal returns, and weaker future accounting performance. A double-sort portfolio placing firms simultaneously in the highest tercile of the AI Narrative Gap and highest tercile of analyst optimism earns a mean return 22.8 percentage points below that of the lowest tercile on both dimensions (t = −5.10). The return reduction in the AI Narrative Gap coefficient is attenuated but not eliminated after controlling for optimism, consistent with a partial expectation-distortion channel. Collectively, these results establish the AI Narrative Gap as a cross-sectionally informative firm characteristic that captures the credibility of a firm’s AI strategic identity, with systematic implications for analyst expectations and asset prices.
  • 详情 Beyond Prompting: An Autonomous Framework for Systematic Factor Investing via Agentic AI
    This paper develops an autonomous framework for systematic factor investing via agentic AI. Rather than relying on sequential manual prompts, our approach operationalizes the model as a self-directed engine that endogenously formulates interpretable trading signals. To mitigate data snooping biases, this closed-loop system imposes strict empirical discipline through out-of-sample validation and economic rationale requirements. Applying this methodology to the U.S. equity market, we document that long-short portfolios formed on the simple linear combination of signals deliver an annualized Sharpe ratio of 2.75 and a return of 54.81%. Finally, our empirics demonstrate that self-evolving AI offers a scalable and interpretable paradigm.
  • 详情 Autonomous Market Intelligence: Agentic AI Nowcasting Predicts Stock Returns
    Can fully agentic AI nowcast stock returns? We deploy a state-of-the-art Large Language Model to evaluate the attractiveness of each Russell 1000 stock each trading day, starting in April 2025 when AI web interfaces enabled real-time search. Our data contribution is unique along three dimensions. First, the nowcasting framework is completely out-of-sample and free of look-ahead bias by construction: predictions are collected at the current edge of time, ensuring the AI has no knowledge of future outcomes. Second, this temporal design is irreproducible once the information environment passes. Third, our framework is fully agentic: we do not feed the model curated news or disclosures; it autonomously searches the web, filters sources, and synthesises information into quantitative predictions. We find that AI possesses genuine stock-selection ability, but that its predictive power is concentrated in identifying future winners. A daily value-weighted portfolio of the 20 highestranked stocks earns a Fama-French five-factor plus momentum alpha of 19.4 basis points and an annualised Sharpe ratio of 2.68 over April 2025–March 2026. The same portfolio accumulates roughly 49.0% cumulative return, versus 21.2% for the Russell 1000 benchmark. The strategy is economically implementable: the average bid-ask spread of the daily Top-20 portfolio is 1.79 basis points, less than 10% of gross daily alpha. However, the signal remains asymmetric. Bottom-ranked portfolios generally exhibit alphas close to zero, while the strongest predictive content sits in the extreme top ranks. Delayed-entry tests further show that predictability does not vanish after a single day; rather, the signal remains positive over a broad window of subsequent entry dates, consistent with slow information diffusion rather than a fleeting overnight anomaly.
  • 详情 Technological Momentum in China: Large Language Model Meets Simple Classifications
    This study applies large language models (LLMs) to measure technological links and examines its predictive power in the Chinese stock market. Using the BAAI General Embedding (BGE) model, we extract semantic information from patent textual data to construct the technological momentum measure. As a comparison, the measure based on traditional International Patent Classification (IPC) is also considered. Empirical analysis shows that both measures significantly predict stock returns and they capture complementary dimensions of technological links. Further investigation through stratified analysis reveals the critical role of investor inattention in explaining their differential performance: in stocks with low investor inattention, IPC-based measure loses its predictive power while BGE-based measure remains significant, indicating that straightforward information is fully priced in while complex semantic relationships require greater cognitive processing; in stocks with high investor inattention, both measures exhibit predictability, with BGE-based measure showing stronger effects. These findings support behavioral finance theories suggesting that complex information diffuses more slowly in markets, especially under significant cognitive constraints, and demonstrate LLMs’ advantage in uncovering subtle technological connections that traditional methods overlook.
  • 详情 A multifactor model using large language models and investor sentiment from photos and news: new evidence from China
    This study introduces an innovative approach for constructing multimodal investor sentiment indices and explores their varying impacts on stock market returns. We employ the RoBERTa model to quantify text-based sentiment, the Google Inception(v3) model for image-based sentiment measurement, and a multimodal semantic correlation fusion model to comprehensively consider the interplay between textual and visual sentiment features. These sentiment indices are further categorised into industry-specific investor sentiment and market-wide investor sentiment, enabling separate analyses of their effects on stock markets. Furthermore, we leverage these indices to build a multifactor stock selection model and timing strategies. Our research findings demonstrate that multimodal sentiment analysis yields superior predictive accuracy. Industry-specific investor sentiment exerts bidirectional positive influences on stock market returns, whereas market-wide investor sentiment indices exhibit unidirectional impacts. Integrating industry-specific investor sentiment into our multifactor stock selection model effectively enhances portfolio returns. Furthermore, combining market-wide investor sentiment with timing strategy optimisation further augments this advantage.
  • 详情 Large Language Models and Return Prediction in China
    We examine whether large language models (LLMs) can extract contextualized representation of Chinese news articles and predict stock returns. The LLMs we examine include BERT, RoBERTa, FinBERT, Baichuan, ChatGLM and their ensemble model. We find that tones and return forecasts extracted by LLMs from news significantly predict future returns. The equal- and value-weighted long minus short portfolios yield annualized returns of 90% and 69% on average for the ensemble model. Given that these news articles are public information, the predictive power lasts about two days. More interestingly, the signals extracted by LLMs contain information about firm fundamentals, and can predict the aggressiveness of future trades. The predictive power is noticeably stronger for firms with less efficient information environment, such as firms with lower market cap, shorting volume, institutional and state ownership. These results suggest that LLMs are helpful in capturing under-processed information in public news, for firms with less efficient information environment, and thus contribute to overall market efficiency.
  • 详情 Burden of Improvement: When Reputation Creates Capital Strain in Insurance
    A strong reputation is a cornerstone of corporate finance theory, widely believed to relax financial constraints and lower capital costs. We challenge this view by identifying an ‘reputation paradox’: under modern risk-sensitive regulation, for firms with long-term liabilities, a better reputation may paradoxically increase capital strain. We argue that the improvement of firm’s reputation alters customer behavior , , which extends liability duration and amplifies measured risk. By using the life insurance industry as an ideal laboratory, we develop an innovative framework that integrates LLMs with actuarial cash flow models, which confirms that the improved reputation increases regulatory capital demands. A comparative analysis across major regulatory regimes—C-ROSS, Solvency II, and RBC—and two insurance products, we further demonstrate that improvements in reputation affect capital requirements unevenly across product types and regulatory frameworks. Our findings challenge the conventional view that reputation uniformly alleviates capital pressure, emphasizing the necessity for insurers to strategically align reputation management with solvency planning.