Skip to main content
Hanabi-1 is a compact transformer for financial time-series prediction. Its training addresses class imbalance and combines direction prediction with regression tasks. The model uses:
  • Transformer layers with multi-head attention
  • Shared hidden-state representations
  • Multiple specialized predictive pathways for direction, volatility, price change, and spread
  • Batch normalization to stabilize training
  • Focal loss implementation to address inherent class imbalance
The compact architecture supports edge deployment for real-time decisions.

Mathematical Foundations: Functions and Formulas

Positional Encoding

Sinusoidal positional encoding represents each position in a sequence: PE(pos,2i)=sin⁡(pos100002i/dmodel)PE(pos,2i) = \sin\left(\frac{pos}{10000^{2i/d_{model}}}\right) PE(pos,2i+1)=cos⁡(pos100002i/dmodel)PE(pos,2i+1) = \cos\left(\frac{pos}{10000^{2i/d_{model}}}\right) Where pos is the position within the sequence and i is the dimension index.

Focal Loss for Direction Prediction

Focal loss adjusts the contribution of examples during direction-prediction training: FL(pt)=−(1−pt)γlog⁡(pt)FL(p_t) = -(1-p_t)^\gamma \log(p_t) Where p_t is the model’s estimated probability for the correct class and γ is the focusing parameter. Increasing γ reduces the contribution of examples for which the model already assigns a high probability to the correct class.

Confidence Calibration

Hanabi-1 computes a confidence score from the distance between the predicted probability and the decision threshold: Confidence=2⋅∣p−threshold∣\text{Confidence} = 2 \cdot |p - \text{threshold}| Where p is the predicted probability and threshold is the decision boundary. Users can filter predictions by this score. Chart Image Confidence vs Accuracy The plotted evaluation groups predictions by confidence. The reported accuracy is nearly 100% in the “High” group and close to chance in the “Very Low” group.

Training Dynamics and Balanced Validation

The validation score penalizes imbalanced predictions, including collapse to a single predicted class: ValScore=F1+0.5⋅Accuracy+0.5⋅PRbalance−0.1⋅Loss−Balancepenalty\text{ValScore} = F1 + 0.5 \cdot \text{Accuracy} + 0.5 \cdot PR_{\text{balance}} - 0.1 \cdot \text{Loss} - \text{Balance}_{\text{penalty}} Where PRbalancePR_{\text{balance}} is the precision-recall balance metric: PRbalance=min⁡(Precision,Recall)max⁡(Precision,Recall)PR_{\text{balance}} = \frac{\min(\text{Precision}, \text{Recall})}{\max(\text{Precision}, \text{Recall})} And Balancepenalty\text{Balance}_{\text{penalty}} applies severe penalties for extreme prediction distributions:
Validation uses this score to account for accuracy and prediction balance: score Image Training dynamics The plot records training and validation changes over the displayed run.

Model Architecture Details

The architecture combines temporal aggregation with task-specific prediction pathways:
  • Feature differentiation through multiple temporal aggregations:
    • Last hidden state capture (most recent information)
    • Average pooling across the sequence (baseline signal)
    • Attention-weighted aggregation (focused signal)
  • Direction pathway with BatchNorm for stable training:
    • Fully connected layers with BatchNorm1d
    • LeakyReLU activation to keep a nonzero gradient for negative inputs
    • Xavier initialization with small random bias terms
  • Specialized regression pathways:
    • Separate networks for volatility, price change, and spread prediction
    • Reduced complexity compared to the direction pathway
    • Independent optimization focuses training capacity where needed
The prediction tasks share a transformer encoder, so their training updates the same representations.

Prediction Temporal Distribution

prediction Image Direction Probabilities The plot shows directional predictions over time. Green dots mark correct predictions; red dots mark incorrect predictions.

Performance and Future Directions

Reported evaluation metrics:
  • Direction accuracy: 73.9%
  • F1 score: 0.67
  • Balanced predictions: 54.2% positive / 45.8% negative
The documented input-window configurations are:
  • 4-hour window model (w4_h1)
  • 12-hour window model (w12_h1)
These configurations predict market movements for the next hour. The reported evaluation favored the 12-hour window in more volatile conditions. Future developments include:
  • Extending prediction horizons to 4, 12 and 24 hours
  • Implementing adaptive thresholds based on market volatility
  • Adding meta-learning approaches for hyperparameter optimization
  • Integrating on-chain signals for cross-domain pattern recognition