MORPH VALUE RANK
…Waiting for measurementsBEST VALUE
…Speed relative to workload costMORPH WORKLOAD COST
…Per request at the default workloadDoes caching track value and traffic?
Fetching observed endpoint cache rates
VALUE VS CACHE RATE
…Spearman rank correlationSPEED VS CACHE RATE
…Spearman rank correlationTRAFFIC SHARE VS CACHE RATE
…Spearman rank correlationMORPH OBSERVED CACHE RATE
…Cached input tokens / total input tokens+1 means higher cache rates accompany better scores. 0 means no rank association. −1 means the opposite.
Traffic shares use reported requests over 30 minutes.
The speed / cost tradeoff
Higher and further left is better. Green marks Morph.
Where the value is
Fetching the latest 30 minute window
| Rank | Provider / endpoint | Value index ⓘ | Median TPS | Cost / M output | Observed cache | Traffic share | First token | Samples |
|---|---|---|---|---|---|---|---|---|
| Loading live endpoint measurements… | ||||||||
Formula and methodology
The default workload is 8,192 input tokens and 1,024 output tokens, with no cached input. All precisions are included. Endpoints need at least 100 throughput samples to be ranked.
Cost per million output tokens = output price + 8 × input price. Prices are per million tokens, with published discounts already included. Per request cost = (1,024 × output price + 8,192 × input price) / 1,000,000.
Value = median TPS / cost per million output tokens. Value index = 100 × value / best eligible value. Each endpoint is ranked separately, including fast variants and different precisions. Missing prices or measurements stay unranked.
Observed cache rates are separate measurements, used for the correlation panel and chart. The cost assumes uncached input for every provider, which avoids building an automatic cache advantage into the score we are testing.
Correlation is Spearman's correlation between average ranks, with ties handled equally and one vote per eligible endpoint. Endpoints without cache measurements are excluded, not assigned zero. At least three paired endpoints and variation in both measures are required. It describes association across providers, not whether caching caused the difference. Cache rates come from OpenRouter's effective pricing feed, whose aggregation window is unspecified. Throughput uses a rolling 30 minute window.
Reported traffic share = endpoint request count / total request count across visible standard endpoints for the selected model. The denominator includes endpoints outside the rankings. Only valid request counts from the same 30 minute window are included. Missing counts stay unavailable, and a measured zero is zero when total requests are positive. This is a proxy for traffic based on reported requests, not a verified share of all routed attempts. Traffic correlation pairs that share with observed cache rate across eligible endpoints. The cache aggregation window is unspecified and may differ from the request window.
The speed versus cost and value versus cache charts connect Pareto optimal measurements with a smooth frontier curve. An endpoint is optimal when no other endpoint is at least as good on both axes and strictly better on one. The cost view favors higher throughput and lower cost. The cache view favors higher value with less cache use. Curves only connect observations, they are not estimates of performance between endpoints.
The workload estimates cost, it does not predict throughput. This is a market comparison, not OpenRouter's private routing score.