Skip to main content
Weflayr replays your real production prompts through candidate models and compares their outputs with the current model, and prices the swap against your actual usage. Benchmarks run on their own: once Weflayr has seen enough traffic on a feature, it replays that feature in the background and publishes the result here. A feature is the feature_name you stamp on your calls when propagating metadata. Like prompt caching optimisation, it needs capture message content on (see Cost Optimisation).
Open a benchmarked feature to see its result.

Model Benchmark detail

High level view of the results from the different model’s results, and the retrospective impact on your unit economics.
For every replayed request, an LLM judge compares your current model’s answer and the candidate’s answer, and rates the candidate better, same or worse on five scores:

Per-example results

Better understand the benchmarked model behavior.