Model Benchmarking

The Model Benchmarking tries to find cheaper model that performs as well as the one you run today. Weflayr replays your real production prompts through candidate models and compares their outputs with the current model, and prices the swap against your actual usage.

Benchmarks run on their own: once Weflayr has seen enough traffic on a feature, it replays that feature in the background and publishes the result here. A feature is the feature_name you stamp on your calls when propagating metadata. Like prompt caching, it needs capture message content on (see Cost Optimisation).

Open a benchmarked feature from the Optimisation engine front page to see its result.

Model Benchmark detail

High level view of the results from the different model’s results, and the retrospective impact on your unit economics.

Benchmarked examples

Better understand the benchmarked model behavior.