Do Equities and Bitcoin Run on the Same Fuel? A Modality-Controlled Comparison of Price, Sentiment, and Activity Signals for Next-Day Return Prediction
Darsh Gupta
09/10/2026
Machine-learning research on equity return prediction and on cryptocurrency return prediction has developed as two largely disjoint literatures, each with its own datasets, feature conventions, and architectures. This methodological separation makes it impossible to determine whether the two asset classes are driven by the same informational modalities, or whether—as is widely conjectured—cryptocurrencies are disproportionately attention-driven while equities remain anchored to price momentum. A modality, in this paper, simply means a category of input information: how the price has moved, what the news says, or how much real economic activity underlies the asset.
We address this question through a modality-controlled experimental design, in which an identical modeling pipeline is applied to a large-capitalization equity (AAPL) and to Bitcoin (BTC-USD) over a common two-year daily sample (6 July 2024–6 July 2026; n=499 and n=730 usable observations, respectively). Three feature modalities are constructed symmetrically for both assets: price (moving averages, RSI, MACD, realized volatility), sentiment (news headlines scored for positive or negative tone and aggregated to a daily value), and activity (blockchain telemetry—active addresses, transaction count, hash rate—for Bitcoin, and volume-based order-flow proxies for the equity).
A stacked LSTM is trained under a price-only baseline and under three additive ablations. An LSTM, or Long
Short-Term Memory network, is a neural network designed for data that arrives in time order; it processes each day in sequence while carrying forward a memory of earlier days, which lets it learn patterns that unfold over time. Additive fusion refers to the practice of combining several information types into one model by simply adding all of their features together, and an ablation is the procedure of adding or removing one component at a time to isolate how much that component contributes. A LightGBM surrogate with SHAP attribution ranks modality importance, and Granger causality tests assess cross-asset predictive spillover. SHAP is a method that assigns each input feature a share of the credit for a model’s predictions, so that the features can be ranked by how much they actually mattered. A Granger causality test asks whether knowing one series helps predict another beyond what the second series’ own past already explains.
Three findings emerge. First, contrary to the prevailing framing of the multimodal fusion literature, additive fusion degraded out-of-sample performance for both assets: the price-only baseline attained the highest directional accuracy (52.9% for AAPL, 50.5% for BTC-USD) while full fusion attained the lowest (42.6% and 45.6%), with root-mean-square error inflating by 49.1% and 66.9% respectively. Directional accuracy is the share of days on which the model correctly predicted whether the price would rise or fall, so 50% is equivalent to guessing. We emphasize—and quantify—that none of these accuracies is statistically distinguishable from chance at conventional levels given the available test-set sizes, and we argue this null-result framing is itself the appropriate scientific conclusion.
Second, attribution is markedly asset-heterogeneous, but not in the direction the attention-driven hypothesis predicts: momentum and money-flow constructs (MACD, the money-flow proxy, relative volume, RSI) dominate for the equity, whereas the three highest-ranked features for Bitcoin are all on-chain activity measures (hash rate, transaction count, unique active addresses). On-chain data refers to statistics read directly from Bitcoin’s public transaction ledger, a class of information that has no equivalent for a company’s stock. For both assets the two sentiment features rank last, with attribution distributions collapsed near zero—an attribution-side result that is mutually consistent with the ablation-side failure of fusion.
Third, spillover is nonetheless asymmetric: Bitcoin news sentiment Granger-causes next-day equity returns at lags one and two (p=0.018, p=0.034), while the reverse direction is insignificant at every tested lag (p>0.24). Taken together, these results argue that multimodal fusion should be treated as an empirical hypothesis to be tested per asset and per corpus rather than an assumed improvement; that Bitcoin’s daily predictability, to the extent it exists, is grounded in settlement-layer activity rather than in news polarity; and that any sentiment channel between the two markets operates cross-sectionally rather than within-asset.
