
LLM A/B Testing App
An internal Retool tool that scores two LLMs against a named rubric instead of a bare accuracy number — built from evaluation work on Gemini.
Requires a Retool workspace login — the screens below are the actual UI.
Overview
Most model-comparison tools stop at a percentage. A weighted score that says Model A beat Model B by 64 points doesn’t say why — so whoever’s deciding which model to ship has to trust the tool instead of reading the reasoning. This app configures two models on a real, context-dependent prompt, runs both, and scores them against four named criteria instead of one opaque number.
Problem
A generic accuracy score can’t tell the difference between a response that’s grounded in what the person actually said and one that just sounds confident. And a comparison tool built on a placeholder prompt, with a score that doesn’t match its own breakdown, has the exact problem it exists to catch — just in its own UI instead of the model output.
Revision
The first pass read like a demo, not a result — caught in a design review the same way the tool is meant to catch it in a model.
Before

A placeholder prompt (“explain the health benefits of drinking water”), an 87/23 score with no breakdown to justify it, and a generic “Recommended Actions” card — retrain, add guardrails, normalize temperature — that would apply to any model, on any prompt.
After

A real support-escalation prompt where generic advice visibly fails, and a rubric that names why Model A won on every axis instead of asserting it.
What I cut
The generic “Recommended Actions” card. Retrain the model, add safety guardrails, normalize temperature — none of that follows from a drinking-water prompt. It read as a button that dumps advice, not a reviewer explaining a decision.
What shipped
A rubric — grounded, actionable, doesn’t invent, fit to the setting — scored pass or fail per axis, per model, so the weighted score and the reasoning underneath it agree.
System
The rubric is the same shape as the review criteria from evaluation work on Gemini — not “did the model sound right,” but would a human reviewer sign off on this specific response, in this specific situation. That’s the standard the tool holds itself to as well as the models it compares.
Outcome
A working internal tool — the screens above are the app’s actual UI, not a mockup of one.
