Introduced by TypeSafe AI alongside its System One Model Jev in September 2026, RLCD is explicitly framed as distinct from RLHF (which optimizes for human-preferred text) and RLVR (which optimizes for verifiably correct outputs like passing test cases). Rather than rewarding a specific answer, RLCD rewards calibration across the whole distribution of a model's predictions, the same property statisticians look for in a well-calibrated weather forecaster. As of its introduction, TypeSafe had not published an architecture paper, so RLCD's training mechanics beyond this framing remain undisclosed.