TrainRL · Post-training research

From learning signal
to useful behavior.

TrainRL is Vikram Kharvi's independent research log on post-training language models. The work is released one technical blog at a time so each finding has room to be read, tested, and discussed.

Current release / Blog 01

Post-training is a stack.

Start with the end-to-end map: assistant contracts, supervised fine-tuning, preference data, reward models, verifiable RL, PPO, DPO, tools, memory, security boundaries, and independent evaluation.

Read the foundational guide

Public research / 01 blog

Available now

The remaining field notes are preserved off the public site and will be released gradually.

  1. 01
    Blog 01 · 13 Aug 202615 min read

    Post-training Is a Stack

    The full map of supervised tuning, synthetic data, preferences, RL, safety, agent trajectories, distillation, and continual improvement—with guidance on when each method fits.

    Read

Across Art of Cyber AI

Follow the model lifecycle

The post-training research is active. Pre-training and inference are coming soon.