Indonesian Coal Forecasting Research

Do external market signals and more complex models improve next-day return forecasts for Indonesian coal equities?

Across ADRO, PTBA, and ITMG, neither added external variables nor higher model complexity produced a consistent advantage during the frozen evaluation period over a zero-return baseline.

Why I investigated this

I started with a broader curiosity about Indonesian financial markets. Coal gave me a constrained test case. Several large listed companies share a sector, while market conditions outside their own price histories can affect them.

In 2025, Indonesia produced 790 million tonnes of coal and exported 514 million tonnes, or 65.1% of output, according to the Ministry of Energy and Mineral Resources [13]. That export exposure made coal a useful setting for asking whether external market information appears in listed-equity returns.

That made a practical question possible. If these companies share external pressures, should movements in USD/IDR or peer coal equities help forecast their next trading-day returns?

Prior Indonesian studies made exchange-rate exposure reasonable to test [1] [2]. They did not prove that the variables would improve prediction.

Equities
3 IDX coal stocks
External signal
USD/IDR

The question became harder

My first version was narrower: one stock, a GRU, and USD/IDR as an external signal. When that signal did not clearly help, the assumption behind the model became more interesting than tuning the model.

First question

Can I forecast ADRO with a GRU?

Harder question

Can extra information or complexity survive a stricter experiment?

Questions worth testing

  1. Does USD/IDR add information beyond a stock's own history?
  2. Do peer Indonesian coal equities add useful signal?
  3. Does combining foreign-exchange and peer information help?
  4. Do sequential deep-learning models outperform simpler baselines?
  5. Does a Transformer improve enough over a GRU to justify its extra complexity?

What the data looked like

Series
ADRO.JK · PTBA.JK · ITMG.JK · USD/IDR
Coverage
Daily data · 2016 to 2026
Target
Next observed trading-day adjusted-close log return

Finding 01

Large moves surrounded near-zero means

The line spans the observed minimum and maximum daily return. The white mark is the mean.

Observed daily return ranges and means ADRO ranges from negative 0.2829 to positive 0.1772 with mean 0.00122. PTBA ranges from negative 0.1893 to positive 0.1924 with mean 0.00094. ITMG ranges from negative 0.1052 to positive 0.1815 with mean 0.00114. -0.30 0 +0.20 ADRO -0.2829 +0.1772 PTBA -0.1893 +0.1924 ITMG -0.1052 +0.1815
Means were 0.00122, 0.00094, and 0.00114. A zero-return forecast is therefore a serious baseline, not a token comparison.

Finding 02

The return series had heavy tails

Excess kurtosis was well above zero for all three equities.

Excess kurtosis by equity ADRO excess kurtosis is 7.45, PTBA is 7.80, and ITMG is 3.69. 0 4 8 ADRO 7.45 PTBA 7.80 ITMG 3.69
A Jarque-Bera test rejects normality for all three series. Rolling volatility also changes through time, so the feature set records 5-day and 20-day volatility.

Lagged relationships were weak. The EDA found no strong persistent lead-lag pattern with the external series. ADRO also shows a visible shift around the late-2024 ADRO/AADI restructuring. No formal break test was run, so the study carries this as a limitation rather than claiming a break.

Designing a result I could defend

Forecast returns, not raw prices
Raw prices are persistent. Next-day log returns remove the easy shortcut of predicting something close to today's price.
Use a zero-return baseline
Complexity must beat a forecast of no next-day movement before it earns a claim of added value.
Keep future information out
External observations receive a one-observation availability delay, then backward-only as-of alignment.
Fit preprocessing on past data
Each fold fits its scalers on the training period. Validation and evaluation rows are transform-only.
Preserve time order
Three expanding windows replace random train-test splits, following time-series evaluation guidance [6] [7] [8].
Freeze the later period
Candidate selection uses the development folds. Frozen families refit on pre-2023 data before one later evaluation.
Compare like with like
Feature groups are compared inside the same family. All families use at most the previous 20 trading days.
Freeze one candidate per family
Mean fold RMSE selects one XGBoost, GRU, and Transformer candidate. Families do not compete during selection.

Temporal validation

Each decision stops before its evidence

Three expanding development folds and one frozen evaluation period Fold one trains from 2016 through 2019 and validates on 2020. Fold two trains through 2020 and validates on 2021. Fold three trains through 2021 and validates on 2022. Decisions are frozen, then evaluated from 2023 through April 2026. 2016 2018 2020 2022 2024 2026 Fold 1 train validate 2020 Fold 2 validate 2021 Fold 3 validate 2022 Frozen evaluate 2023 to 2026 select on folds freeze decisions
Intervals are half-open. The frozen period is procedural, not pristine: earlier project iterations had already inspected parts of 2023 to 2026.

Testing complexity

This model ladder tests whether each increase in complexity produces a repeatable reduction in forecast error.

  1. Zero-return baseline

    Sets the minimum useful performance with a forecast of 0.0.

  2. XGBoost

    Tests bounded nonlinear relationships in engineered features [9].

  3. GRU

    Tests whether short sequential structure adds value [10].

  4. Transformer

    Tests extra sequence-model complexity, with counterevidence in view [11] [12].

Testing information

E0
Own historyThe target's return history, represented as engineered summaries or raw sequences.
E1
E0 + USD/IDRLagged FX information with the conservative availability rule.
E2
E0 + peersLagged returns from the other coal equities.
E3
E0 + USD/IDR + peersThe two external signal sets combined.

E1 follows Indonesian evidence on exchange-rate exposure [1] [2].

E2 is exploratory. Industry information diffusion makes peer returns plausible, but the source evidence is monthly US data, not daily Indonesian coal data [3] [4] [5].

I limited the external features to two ideas with a clear research rationale: exchange-rate exposure and peer-market information. The ablation tests whether either one actually improves forecasting.

What survived evaluation

Some complex models produced small improvements on individual equities. Those gains did not repeat consistently across stocks or feature groups.

Model comparison · E0 only

Difference in final RMSE versus zero return

Negative is lower error. Positive is worse. The baseline sits at 0% for each equity.

E0 model RMSE difference from the zero-return baseline For ADRO, XGBoost is 2.182 percent worse, GRU is 0.018 percent better, and Transformer is 0.631 percent worse. For PTBA, XGBoost is 2.169 percent worse, GRU is 0.165 percent better, and Transformer is 0.585 percent worse. For ITMG, XGBoost is 0.928 percent worse, GRU is 0.312 percent worse, and Transformer is 0.947 percent worse. -0.25% 0% +1% +2% ADRO XGBoost +2.182% GRU -0.018% Transformer +0.631% PTBA XGBoost +2.169% GRU -0.165% Transformer +0.585% ITMG XGBoost +0.928% GRU +0.312% Transformer +0.947%

ADRO · Δ RMSE vs naive

XGBoost+2.182%
GRU-0.018%
Transformer+0.631%

PTBA · Δ RMSE vs naive

XGBoost+2.169%
GRU-0.165%
Transformer+0.585%

ITMG · Δ RMSE vs naive

XGBoost+0.928%
GRU+0.312%
Transformer+0.947%
  • Lower RMSE
  • Naive baseline
  • Higher RMSE
Frozen families were refit on pre-2023 development data, then scored once on the 787-row forecasting evaluation set. This is a pipeline comparison under a shared 20-day information horizon, not a pure architecture comparison.
Exact final RMSE · E0 own-history features
Equity Naive XGBoost GRU Transformer
ADRO 0.027035 0.027625 0.027030 0.027205
PTBA 0.019960 0.020393 0.019927 0.020077
ITMG 0.017460 0.017622 0.017515 0.017626

Feature ablation

Development wins sometimes changed direction in the frozen period.

ADRO's GRU with peer features improved mean validation RMSE by 0.000290 versus E0. In frozen evaluation it was worse by 0.000064.

ADRO · GRU · E2 versus E0

RMSE delta. Negative is lower error.

ADRO GRU peer-feature RMSE delta changed direction Mean development validation delta was negative 0.000290. Frozen evaluation delta was positive 0.000064. -0.00032 0 +0.00010 Validation -0.000290 Evaluation +0.000064 lower error higher error
Final RMSE versus E0 inside each family · ↓ lower · ↑ higher
Family / group ADRO PTBA ITMG Lower error
XGBoost · E1↓↓↑2/3
XGBoost · E2↓↑↑1/3
XGBoost · E3↓↓↑2/3
GRU · E1↓↑↑1/3
GRU · E2↑↑↑0/3
GRU · E3↑↑↑0/3
Transformer · E1↑↑↑0/3
Transformer · E2↓↓↑2/3
Transformer · E3↓↓↑2/3

No external feature group improved all three equities within any model family.

How to read the result

RMSE measures error size, not economic value. The small differences do not support a trading claim, a causal claim, or a general claim about all Indonesian equities. They describe three folds and one fixed evaluation period.

10/12The GRU had the lowest mean development-fold RMSE in 10 of 12 equity × feature-group comparisons. That looked promising. It was not enough.

What the evidence changed in my thinking

A complex model can look better during development and still offer little evidence of a repeatable advantage. I would rather keep that negative result than turn a fragile improvement into a forecasting claim.

Additional analysis

I also estimated one-day 95% historical VaR from returns before 2023, then counted breaches in the frozen period. It is a descriptive risk measure, separate from model selection and the forecasting conclusion.

ADRO
0.042849IDR 41,944.32 loss on IDR 1,000,000
26/788 breaches
PTBA
0.038726IDR 37,986.16 loss on IDR 1,000,000
18/788 breaches
ITMG
0.039933IDR 39,146.02 loss on IDR 1,000,000
20/788 breaches

The estimate is unconditional and uses adjusted-close log returns. It uses no forecast or residual, drives no selection, and does not guarantee future coverage.

Limits and the next test

Frozen, not untouched

The later period is not a pristine test set

Earlier iterations inspected parts of 2023 to 2026 before the current method was final. Treat it as honest evidence for the frozen protocol, not proof against all future peeking.

Possible regime change

ADRO changes visibly in late 2024

The ADRO/AADI restructuring creates a plausible structural change. The experiment does not model corporate-action interpretation or stable pre-event and post-event regimes.

Omitted by design

Coal-price information is incomplete

The repository has no historical release-dated HBA series with enough publication timing detail for leak-safe alignment. Backfilling a later-known monthly value would leak information.

  • Reserve a genuinely unseen future period
  • Model corporate-event regimes explicitly
  • Add release-dated HBA or another coal benchmark
  • Expand the cross-sectional sample

Research reference

These sources explain why a variable, comparison, or model entered the experiment. They do not validate this study's result.

  1. A. R. Putra and R. Robiyanto, "The Effect of Commodity Price Changes and USD/IDR Exchange Rate on Indonesian Mining Companies' Stock Return," Jurnal Keuangan dan Perbankan, vol. 23, no. 1, pp. 97-108, 2019. doi:10.26905/jkdp.v23i1.2084.
  2. A. A. Komara, B. M. Sinaga, and T. Andati, "The Impact of Changes in External and Internal Factors on Financial Performance and Stock Returns of Coal Companies," Jurnal Aplikasi Bisnis dan Manajemen, vol. 5, no. 3, p. 513, 2019. doi:10.17358/jabm.5.3.513.
  3. T. J. Moskowitz and M. Grinblatt, "Do Industries Explain Momentum?" The Journal of Finance, vol. 54, no. 4, pp. 1249-1290, 1999. doi:10.1111/0022-1082.00146.
  4. K. Hou, "Industry Information Diffusion and the Lead-Lag Effect in Stock Returns," The Review of Financial Studies, vol. 20, no. 4, pp. 1113-1138, 2007.
  5. U. Ali and D. Hirshleifer, "Shared Analyst Coverage: Unifying Momentum Spillover Effects," Journal of Financial Economics, vol. 136, no. 3, 2020.
  6. L. J. Tashman, "Out-of-Sample Tests of Forecasting Accuracy: An Analysis and Review," International Journal of Forecasting, vol. 16, pp. 437-450, 2000. doi:10.1016/S0169-2070(00)00065-0.
  7. C. Bergmeir, R. J. Hyndman, and B. Koo, "A Note on the Validity of Cross-Validation for Evaluating Autoregressive Time Series Prediction," Computational Statistics and Data Analysis, vol. 120, pp. 70-83, 2018. doi:10.1016/j.csda.2017.11.003.
  8. V. Cerqueira, L. Torgo, and I. Mozetič, "Evaluating Time Series Forecasting Models: An Empirical Study on Performance Estimation Methods," Machine Learning, vol. 109, pp. 1997-2028, 2020. doi:10.1007/s10994-020-05910-7.
  9. T. Chen and C. Guestrin, "XGBoost: A Scalable Tree Boosting System," Proc. 22nd ACM SIGKDD, pp. 785-794, 2016. doi:10.1145/2939672.2939785.
  10. K. Cho et al., "Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation," Proc. EMNLP, pp. 1724-1734, 2014. doi:10.3115/v1/D14-1179.
  11. A. Vaswani et al., "Attention Is All You Need," Advances in Neural Information Processing Systems 30, 2017. proceedings.neurips.cc.
  12. A. Zeng, M. Chen, L. Zhang, and Q. Xu, "Are Transformers Effective for Time Series Forecasting?" Proc. AAAI, vol. 37, pp. 11121-11128, 2023. doi:10.1609/aaai.v37i9.26317.
  13. Kementerian Energi dan Sumber Daya Mineral Republik Indonesia, "Capaian Positif Tahun 2025, Negara Hadir Penuhi Kebutuhan Energi Masyarakat," Press Release No. 002.Pers/04/SJI/2026, Jan. 9, 2026. esdm.go.id.

Repository

The repository includes temporal splits, leakage checks, deterministic experiments, and CI tests.

View GitHub Repository