Learning

  • 详情 Determinants of Firm Survival Using Machine Learning: Evidence from the Pearl River Delta, China
    Firm survival, as a key indicator of regional economic resilience, has gained increasing attention in the context of global economic uncertainty and the deep adjustments in industrial structures. In the Pearl River Delta (PRD), a core region of China’s Guangdong-Hong Kong-Macau Greater Bay Area, the characteristics of firm life cycles are crucial for understanding spatial development inequalities and institutional effects in emerging economies. This study focuses on firms in the PRD, using full life-cycle data from registration, operation, to deregistration. An XGBoost regression model is employed, incorporating the SHAP explanation algorithm, to systematically analyze the main factors influencing firm survival. The results show that: (1) Firm establishment time is the primary factor influencing survival, with significant “survival threshold” and “growth leap” effects—mature firms exhibit a distinct survival advantage; (2) Among spatial structure variables, moderate industry specialization and diversity enhance firm survival rates, while excessive concentration and diversification show diminishing or negative returns, reflecting an “ecological threshold” effect; (3) External shocks have a significant suppressive impact on startups, while policy support and capital size show limited explanatory power; (4) The ownership structure, especially state-holding, has a positive moderating effect on firm survival in specific contexts, indicating that private enterprises’ flexibility and adaptability can compensate for institutional gaps. This study offers insights into the spatial heterogeneity of firm survival mechanisms, providing quantitative evidence for regional economic policy adjustments and firm resilience-building. Recommendations include promoting differentiated policies, optimizing industrial ecological spatial layouts, and piloting systems to enhance survival resilience, especially for mixed-ownership firms.
  • 详情 News Sentiment and Overnight Return Prediction: Aid or Redundancy? Evidence from a Large Language Model
    We investigate whether overnight news sentiment adds predictive value for overnight returns. We focus on the CSI300 Index, whose ETFs are widely held by Chinese retail investors. Sentiment indi-cators are constructed from minute-level overnight news using a fine-tuned RoBERTa model. These indicators are combined with market-based variables to predict overnight returns via regression and machine learning. Results show that while the sentiment alone has predictive value, its incremental contribution disappears once the A50 overnight return is included.
  • 详情 Missing Financial Data in Chinese Market
    This paper studies missing firm characteristics in the Chinese stock market and their implications for empirical asset pricing. Relative to the U.S. market, missing firm characteristics in China remain underexplored despite substantial differences in data availability and disclosure environments. Using a dataset of 106 firm characteristics from 1992 to 2021, we document a pronounced cliff-shaped pattern in missingness, with missing rates falling sharply after 2000. We then compare expectation-maximization (EM) and mean imputation (MN) in both univariate characteristic-sorted portfolios and machine-learning applications that combine many predictors. Results indicate that, in univariate analysis, the two methods produce very similar return spreads because they assign largely the same stocks to the extreme deciles. In machine-learning applications, however, EM-imputed data generally produce better-performing prediction-sorted portfolios than mean-imputed data. These findings provide new evidence on missing firm characteristics in a major emerging market and highlight the importance of imputation choices in machine-learning asset-pricing applications.
  • 详情 A Study of the Microdynamics of Early Childhood Learning
    This paper investigates the weekly evolution of child skills as measured by unique data from a widely-emulated early childhood home-visiting program developed in Jamaica, adapted to rural China, and applied in different versions worldwide. The design of the study avoids problems of endogeneity of inputs and lack of truly comparable measures of skills across children that plague previous econometric studies of child development. Skills that are nominally classified as the same, in fact, do not appear to share a common unit scale across levels. They are produced by skill-specific, lifecycle-stage-specific technologies. We formulate and estimate a new dynamic stochastic skill production model for multiple skills that is consistent with the evidence. We quantify the dynamics of early life learning. The model explains the “fadeout” of measures of learning by the emergence of new skills not properly measured. We investigate the role of ability in learning. We find important differences in learning patterns between boys and girls.
  • 详情 Automated Trading System for Straddle-Option Based on Deep Q-Learning
    Straddle Option is a financial trading tool that explores volatility premiums in high-volatility markets without predicting price direction. Although deep reinforcement learning has emerged as a powerful approach to trading automation in financial markets, existing work mostly focused on predicting price trends and making trading decisions by combining multidimensional datasets like blogs and videos, which led to high computational costs and unstable performance in high-volatility markets. To tackle this challenge, we develop automated straddle option trading based on reinforcement learning and attention mechanisms to handle unpredictability in high-volatility markets. Firstly, we leverage the attention mechanisms in Transformer DDQN through both self-attention with time series data and channel attention with multi-cycle information. Secondly, a novel reward function considering excess earnings is designed to focus on long-term profits and neglect short-term losses over a stop line. Thirdly, we identify the resistance levels to provide reference information when great uncertainty in price movements occurs with intensified battle between the buyers and sellers. Through extensive experiments on the Chinese stock, Brent crude oil, and Bitcoin markets, our attention-based Transformer-DDQN model exhibits the lowest maximum drawdown across all markets, and outperforms other models by 92.5% in terms of the average return excluding the crude oil market due to relatively low fluctuation.
  • 详情 Learning, Price Discovery, and Macroeconomic Announcements
    We examine price discovery after irregularly scheduled macroeconomic announce-ments. Exploiting time variation in Chinese macro announcements released outside regular trading hours, this paper isolates the role of elapsed non-trading time in facilitating investor learning and price discovery upon market reopening. We show that longer non-trading intervals generate more efficient post-announcement price discovery, reduce information asymmetry, and diminish subsequent intraday return reversals. The mechanism operates through enhanced retail investor learning: during non-trading hours, retail investors actively acquire information, subsequently trade more aggressively, earn higher profits, and face reduced informational disadvantages at market opening. Our findings highlight that retail investor learning during non-trading hours levels the informational playing field among heterogeneous investors and improves price quality around irregularly timed macroeconomic announcements. These results have broader implications for emerging markets, which similarly feature irregular announcement timing and large populations of uninformed retail investors.
  • 详情 Luck in the Marketplace: Auspicious Timing and Financial Decision-Making
    We study the role of superstition in China’s peer-to-peer lending market by ex-amining whether lenders time their bids according to “lucky hours” from the Chinese farmer’s calendar. Loans funded during lucky hours perform better—but only because the platform lists higher-rated loans at those times. This pattern is consistent with a screening mechanism: highly risk-averse lenders place greater value on both true risk reductions and auspicious-day signals, so the platform maximizes surplus by bundling the two—listing low-risk loans on auspicious days. Moreover, listing safer loans at lucky hours can further boost proffts because biased beliefs decay more slowly under asymmetric (bad-news-heavy) learning.
  • 详情 Can Artificial Intelligence Reduce Corporate Stock Price Crash Risk in China?
    This study examines the effect of artificial intelligence (AI) adoption on stock price crash risk using panel data from Chinese A-share listed firms from 2001 to 2022. We find that higher levels of AI application significantly reduce crash risk, primarily by enhancing information transparency, easing financial constraints, and promoting innovation. Notably, AI improves transparency within supply chains by reducing information asymmetry between upstream and downstream firms, thereby enhancing information flow and reducing market frictions. Among AI types, machine learning proves most effective in lowering crash risk due to its data-processing and forecasting capabilities, while natural language processing and computer vision show weaker effects. The impact of AI is particularly pronounced in non-government-regulated industries and high-tech firms. Moreover, its risk-mitigating effect becomes increasingly significant over time. These results are robust to instrumental variable estimation and staggered difference-in-differences (DID) designs. These findings highlight the strategic role of AI in risk management and offer practical implications for firms and policymakers aiming to enhance transparency, financial resilience, and long-term value creation.
  • 详情 Emotions and Fund Flows: Evidence from Managers' Live Streams
    Do investors respond to what fund managers say, or how they look saying it? Using 2,000 live-streamed sessions by Chinese ETF managers and multimodal machine learning, we show that managers’ facial expressions, not their words, drive fund flows. A one-standard-deviation increase in positive facial affect raises next-day flows by 0.17pp (260% of mean). Vocal tone shows weak effects; textual sentiment shows none. Critically, facial expressions predict flows but not returns, indicating pure persuasion rather than information transmission. Effects strengthen when investors are emotionally vulnerable (down markets, retail-heavy funds) and persist 2-3 weeks before dissipating. Our findings challenge the emphasis on textual disclosure in finance and raise questions about investor protection as video communication proliferates.
  • 详情 Reinforcement Learning and Trading on Noise in Limit Order Markets
    This paper introduces reinforcement learning to examine the effect of trading on noise in a dynamic limit order market equilibrium. It shows that intensive noise liquidity provision (consumption) increases speculators' liquidity consumption (provision), improving (reducing) market liquidity. Channeled by uninformed chasing and informed aggressive liquidity provision, the increasing noise liquidity provision and consumption, respectively, improve price efficiency, generating a U-shaped price efficiency to the noise trading uncertainty on liquidity provision and consumption. Associated with a hump-shaped (U-shaped) profitability for the informed (uninformed) at a U-shaped noise trading cost in the noise trading uncertainty, this implies that, at increasing noise trading cost, intensive noise liquidity provision improves market liquidity, price efficiency, order profitability of informed traders, and reduces the loss, even makes profit, for uninformed traders.