Current release / Blog 01
Post-training is a stack.
Start with the end-to-end map: assistant contracts, supervised fine-tuning, preference data, reward models, verifiable RL, PPO, DPO, tools, memory, security boundaries, and independent evaluation.
Read the foundational guidePublic research / 01 blog
Available now
The remaining field notes are preserved off the public site and will be released gradually.
Across Art of Cyber AI
Follow the model lifecycle
The post-training research is active. Pre-training and inference are coming soon.