back
MIDI-GPT
PyTorch · miditok · DVC · MLflow · Optuna · music21 · ONNX · FastAPI
status:active
ref:ML-02
size:—

Decoder-only transformer implemented from scratch in PyTorch for conditioned symbolic music generation. Every component written by hand: multi-head self-attention, causal masking, positional embeddings, residual connections, layer normalisation, dropout. Not a HuggingFace wrapper.

The current model is 6.7M parameters. Broad pretraining on 230M tokens from GiantMIDI, then fine-tuned on MAESTRO for a focus on classical piano performance. REMI tokenisation via miditok with a 298-token vocabulary. Achieves perplexity 4.7 and produces recognisably musical output.

Conditioning is multi-dimensional: composer, period, form, and key, derived from musicological analysis of 1,276 MAESTRO pieces via title regex and music21. Coverage is 89.5% of forms and 54.5% of keys across the dataset.

The data pipeline runs end to end through DVC: REMI tokenisation, bar-boundary chunking with 50% overlap, conditioning token prepending, training with cosine LR decay and gradient checkpointing. 40+ experiments tracked with MLflow. Hyperparameter search via Optuna TPE.

Validated a pretrain/fine-tune strategy against published benchmarks (Lehmkuhl et al. 2025), confirming Aria-MIDI pretrain then MAESTRO fine-tune outperforms MAESTRO-only training at equivalent parameter count.

Next phase is a 300M parameter cloud training run on an H200 with RoPE and sliding-window attention on a roughly 5B token Aria-MIDI deduped corpus. The pipeline is designed and validated against published benchmarks but not yet executed.

Serving via ONNX export to a FastAPI endpoint. Client-side Tone.js audio synthesis for browser playback.

— —