Do external market signals and more complex models improve next-day return forecasts for Indonesian coal equities?
Across ADRO, PTBA, and ITMG, neither added external variables nor higher model complexity produced a consistent advantage during the frozen evaluation period over a zero-return baseline.
Independent research
Time-series forecasting
Indonesia
Python / PyTorch / XGBoost
Why I investigated this
I started with a broader curiosity about Indonesian financial markets. Coal gave me a constrained test case. Several large listed companies share a sector, while market conditions outside their own price histories can affect them.
In 2025, Indonesia produced 790 million tonnes of coal and exported 514 million tonnes, or 65.1% of output, according to the Ministry of Energy and Mineral Resources [13]. That export exposure made coal a useful setting for asking whether external market information appears in listed-equity returns.
That made a practical question possible. If these companies share external pressures, should movements in USD/IDR or peer coal equities help forecast their next trading-day returns?
Prior Indonesian studies made exchange-rate exposure reasonable to test [1][2]. They did not prove that the variables would improve prediction.
Equities
3 IDX coal stocks
External signal
USD/IDR
The question became harder
My first version was narrower: one stock, a GRU, and USD/IDR as an external signal. When that signal did not clearly help, the assumption behind the model became more interesting than tuning the model.
First question
Can I forecast ADRO with a GRU?
→
Harder question
Can extra information or complexity survive a stricter experiment?
Questions worth testing
Does USD/IDR add information beyond a stock's own history?
Do peer Indonesian coal equities add useful signal?
Does combining foreign-exchange and peer information help?
Do sequential deep-learning models outperform simpler baselines?
Does a Transformer improve enough over a GRU to justify its extra complexity?
What the data looked like
Series
ADRO.JK · PTBA.JK · ITMG.JK · USD/IDR
Coverage
Daily data · 2016 to 2026
Target
Next observed trading-day adjusted-close log return
Finding 01
Large moves surrounded near-zero means
The line spans the observed minimum and maximum daily return. The white mark is the mean.
Means were 0.00122, 0.00094, and 0.00114. A zero-return forecast is therefore a serious baseline, not a token comparison.
Finding 02
The return series had heavy tails
Excess kurtosis was well above zero for all three equities.
A Jarque-Bera test rejects normality for all three series. Rolling volatility also changes through time, so the feature set records 5-day and 20-day volatility.
Lagged relationships were weak. The EDA found no strong persistent lead-lag pattern with the external series. ADRO also shows a visible shift around the late-2024 ADRO/AADI restructuring. No formal break test was run, so the study carries this as a limitation rather than claiming a break.
Designing a result I could defend
Forecast returns, not raw prices
Raw prices are persistent. Next-day log returns remove the easy shortcut of predicting something close to today's price.
Use a zero-return baseline
Complexity must beat a forecast of no next-day movement before it earns a claim of added value.
Keep future information out
External observations receive a one-observation availability delay, then backward-only as-of alignment.
Fit preprocessing on past data
Each fold fits its scalers on the training period. Validation and evaluation rows are transform-only.
Preserve time order
Three expanding windows replace random train-test splits, following time-series evaluation guidance [6][7][8].
Freeze the later period
Candidate selection uses the development folds. Frozen families refit on pre-2023 data before one later evaluation.
Compare like with like
Feature groups are compared inside the same family. All families use at most the previous 20 trading days.
Freeze one candidate per family
Mean fold RMSE selects one XGBoost, GRU, and Transformer candidate. Families do not compete during selection.
Temporal validation
Each decision stops before its evidence
Intervals are half-open. The frozen period is procedural, not pristine: earlier project iterations had already inspected parts of 2023 to 2026.
Testing complexity
This model ladder tests whether each increase in complexity produces a repeatable reduction in forecast error.
Zero-return baseline
Sets the minimum useful performance with a forecast of 0.0.
XGBoost
Tests bounded nonlinear relationships in engineered features [9].
GRU
Tests whether short sequential structure adds value [10].
Transformer
Tests extra sequence-model complexity, with counterevidence in view [11][12].
Testing information
E0
Own historyThe target's return history, represented as engineered summaries or raw sequences.
E1
E0 + USD/IDRLagged FX information with the conservative availability rule.
E2
E0 + peersLagged returns from the other coal equities.
E3
E0 + USD/IDR + peersThe two external signal sets combined.
E1 follows Indonesian evidence on exchange-rate exposure [1][2].
E2 is exploratory. Industry information diffusion makes peer returns plausible, but the source evidence is monthly US data, not daily Indonesian coal data [3][4][5].
I limited the external features to two ideas with a clear research rationale: exchange-rate exposure and peer-market information. The ablation tests whether either one actually improves forecasting.
What survived evaluation
Some complex models produced small improvements on individual equities. Those gains did not repeat consistently across stocks or feature groups.
Model comparison · E0 only
Difference in final RMSE versus zero return
Negative is lower error. Positive is worse. The baseline sits at 0% for each equity.
ADRO · Δ RMSE vs naive
XGBoost+2.182%
GRU-0.018%
Transformer+0.631%
PTBA · Δ RMSE vs naive
XGBoost+2.169%
GRU-0.165%
Transformer+0.585%
ITMG · Δ RMSE vs naive
XGBoost+0.928%
GRU+0.312%
Transformer+0.947%
Lower RMSE
Naive baseline
Higher RMSE
Frozen families were refit on pre-2023 development data, then scored once on the 787-row forecasting evaluation set. This is a pipeline comparison under a shared 20-day information horizon, not a pure architecture comparison.
Exact final RMSE · E0 own-history features
Equity
Naive
XGBoost
GRU
Transformer
ADRO
0.027035
0.027625
0.027030
0.027205
PTBA
0.019960
0.020393
0.019927
0.020077
ITMG
0.017460
0.017622
0.017515
0.017626
Feature ablation
Development wins sometimes changed direction in the frozen period.
ADRO's GRU with peer features improved mean validation RMSE by 0.000290 versus E0. In frozen evaluation it was worse by 0.000064.
ADRO · GRU · E2 versus E0
RMSE delta. Negative is lower error.
Final RMSE versus E0 inside each family · ↓ lower · ↑ higher
Family / group
ADRO
PTBA
ITMG
Lower error
XGBoost · E1
↓
↓
↑
2/3
XGBoost · E2
↓
↑
↑
1/3
XGBoost · E3
↓
↓
↑
2/3
GRU · E1
↓
↑
↑
1/3
GRU · E2
↑
↑
↑
0/3
GRU · E3
↑
↑
↑
0/3
Transformer · E1
↑
↑
↑
0/3
Transformer · E2
↓
↓
↑
2/3
Transformer · E3
↓
↓
↑
2/3
No external feature group improved all three equities within any model family.
How to read the result
RMSE measures error size, not economic value. The small differences do not support a trading claim, a causal claim, or a general claim about all Indonesian equities. They describe three folds and one fixed evaluation period.
10/12The GRU had the lowest mean development-fold RMSE in 10 of 12 equity × feature-group comparisons. That looked promising. It was not enough.
What the evidence changed in my thinking
A complex model can look better during development and still offer little evidence of a repeatable advantage. I would rather keep that negative result than turn a fragile improvement into a forecasting claim.
Additional analysis
I also estimated one-day 95% historical VaR from returns before 2023, then counted breaches in the frozen period. It is a descriptive risk measure, separate from model selection and the forecasting conclusion.
ADRO
0.042849IDR 41,944.32 loss on IDR 1,000,000 26/788 breaches
PTBA
0.038726IDR 37,986.16 loss on IDR 1,000,000 18/788 breaches
ITMG
0.039933IDR 39,146.02 loss on IDR 1,000,000 20/788 breaches
The estimate is unconditional and uses adjusted-close log returns. It uses no forecast or residual, drives no selection, and does not guarantee future coverage.
Limits and the next test
Frozen, not untouched
The later period is not a pristine test set
Earlier iterations inspected parts of 2023 to 2026 before the current method was final. Treat it as honest evidence for the frozen protocol, not proof against all future peeking.
Possible regime change
ADRO changes visibly in late 2024
The ADRO/AADI restructuring creates a plausible structural change. The experiment does not model corporate-action interpretation or stable pre-event and post-event regimes.
Omitted by design
Coal-price information is incomplete
The repository has no historical release-dated HBA series with enough publication timing detail for leak-safe alignment. Backfilling a later-known monthly value would leak information.
Reserve a genuinely unseen future period
Model corporate-event regimes explicitly
Add release-dated HBA or another coal benchmark
Expand the cross-sectional sample
Research reference
These sources explain why a variable, comparison, or model entered the experiment. They do not validate this study's result.
A. R. Putra and R. Robiyanto, "The Effect of Commodity Price Changes and USD/IDR Exchange Rate on Indonesian Mining Companies' Stock Return," Jurnal Keuangan dan Perbankan, vol. 23, no. 1, pp. 97-108, 2019. doi:10.26905/jkdp.v23i1.2084.
A. A. Komara, B. M. Sinaga, and T. Andati, "The Impact of Changes in External and Internal Factors on Financial Performance and Stock Returns of Coal Companies," Jurnal Aplikasi Bisnis dan Manajemen, vol. 5, no. 3, p. 513, 2019. doi:10.17358/jabm.5.3.513.
T. J. Moskowitz and M. Grinblatt, "Do Industries Explain Momentum?" The Journal of Finance, vol. 54, no. 4, pp. 1249-1290, 1999. doi:10.1111/0022-1082.00146.
K. Hou, "Industry Information Diffusion and the Lead-Lag Effect in Stock Returns," The Review of Financial Studies, vol. 20, no. 4, pp. 1113-1138, 2007.
U. Ali and D. Hirshleifer, "Shared Analyst Coverage: Unifying Momentum Spillover Effects," Journal of Financial Economics, vol. 136, no. 3, 2020.
L. J. Tashman, "Out-of-Sample Tests of Forecasting Accuracy: An Analysis and Review," International Journal of Forecasting, vol. 16, pp. 437-450, 2000. doi:10.1016/S0169-2070(00)00065-0.
C. Bergmeir, R. J. Hyndman, and B. Koo, "A Note on the Validity of Cross-Validation for Evaluating Autoregressive Time Series Prediction," Computational Statistics and Data Analysis, vol. 120, pp. 70-83, 2018. doi:10.1016/j.csda.2017.11.003.
V. Cerqueira, L. Torgo, and I. Mozetič, "Evaluating Time Series Forecasting Models: An Empirical Study on Performance Estimation Methods," Machine Learning, vol. 109, pp. 1997-2028, 2020. doi:10.1007/s10994-020-05910-7.
T. Chen and C. Guestrin, "XGBoost: A Scalable Tree Boosting System," Proc. 22nd ACM SIGKDD, pp. 785-794, 2016. doi:10.1145/2939672.2939785.
K. Cho et al., "Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation," Proc. EMNLP, pp. 1724-1734, 2014. doi:10.3115/v1/D14-1179.
A. Vaswani et al., "Attention Is All You Need," Advances in Neural Information Processing Systems 30, 2017. proceedings.neurips.cc.
A. Zeng, M. Chen, L. Zhang, and Q. Xu, "Are Transformers Effective for Time Series Forecasting?" Proc. AAAI, vol. 37, pp. 11121-11128, 2023. doi:10.1609/aaai.v37i9.26317.
Kementerian Energi dan Sumber Daya Mineral Republik Indonesia, "Capaian Positif Tahun 2025, Negara Hadir Penuhi Kebutuhan Energi Masyarakat," Press Release No. 002.Pers/04/SJI/2026, Jan. 9, 2026. esdm.go.id.
Repository
The repository includes temporal splits, leakage checks, deterministic experiments, and CI tests.