Thesys hosts Generative UI Bench at openui.com/benchmarks and built it partly from OpenUI Cloud's own production traffic. It grades structural validity (does output parse, does every reference resolve, are props missing or out of range, are components invented or orphaned) rather than subjective quality. Thesys used it to score its own OUI-1 model at 71.7%, ahead of Gemma 4 31B (46.7%) but behind Qwen3.8 27B (78.8%) on the same table. No independent publisher or academic group runs or verifies the benchmark, which is a standard caveat for vendor-built evals.