meerada
Models Frontier Labs New Exchange Handshake LLManager Grade board LIVE πŸŽ›οΈ Open LLManager (live) β†’

Every AI model. Measured. Priced. Ranked.

The live model exchange. Every model from every lab β€” graded hourly on tasks a program can verify, priced by the real cost of a done task, tracked the moment it launches. Free, for everyone.
🟒 β€” measured live πŸ“š β€” models tracked πŸ†• β€” launched this month πŸ’± priced by CPAT β€” cost per verified task πŸ”„ refreshed hourly
πŸŽ›οΈ LLManager appOne app for every model, on your own keys. Import your Claude / ChatGPT history and switch models mid-conversation. Free Β· Pro $99/yr. β–Ά Try it in the browserThe live cockpit, no install: run many models on many tasks in parallel, right now. 🀝 Handshake ReportFor companies: a measured migration plan for your real workload β€” where to move, what you save, where quality breaks. πŸ§ͺ Sponsored gradingLabs: get your model on the full battery, daily for 30 days. Results publish either way. ⬇️ Open-source coreThe measuring engine, the CLI and this exchange's data β€” free, on GitHub.

πŸ“š Every model β€” live

Real list prices, context and modalities for every model on the market (public feed, refreshed hourly), with our measured grade and CPAT wherever we have run the battery. Search, filter, sort, and tick up to three to compare side by side.
default: measured models by grade, then the newest launches Β· 50 at a time

How a grade is earned

We give every model the same real, checkable tasks and score only what actually works β€” not opinions.
β€”
πŸ“€
We send a real taskβ€”
πŸ€–
The model answersβ€”
πŸ”
We check it's correctβ€”

What Meerada sells (everything above stays free)

πŸŽ›οΈ LLManager β€” one app for every model, on your keys; import your Claude/ChatGPT history and switch models mid-conversation. Free tier Β· Pro $99/yr.
🀝 Handshake Report β€” a measured migration plan for a company's real workload.
πŸ§ͺ Sponsored grading β€” labs pay for a full battery run; results publish either way.

πŸ›οΈ The labs β€” who makes the models, and what they're worth

The exchange's "issuers": every lab behind the models above, with its parent company or listing, last reported valuation (private) or ticker (public), how many models it fields, its best measured grade and its cheapest verified task. Valuations are last-reported figures with a date and a source β€” not live prices.
Grades attribute to the lab that made the model, wherever it is served β€” e.g. OpenAI's open-weights gpt-oss models measured on Groq's free tier count for OpenAI.
LabParent / listingTickerValuationModelsBest gradeCheapest CPATLatest launch

πŸ’± Outcome Exchange β€” buy a done task, not tokens

This is what makes the arena a real market. Post the job and the quality bar; the exchange routes to the cheapest model that clears it. The price is the live CPAT β€” cost per verified task. As grades move, the clearing price moves. prototype
🧭 The Bourse in 30 seconds
1 Β· Every hour we run the same verifiable tasks (schema, regex, unit-checkable) on every model we can reach, and check the outputs by program β€” not by opinion.
2 Β· Each model gets a grade (verified success rate, with its confidence interval) and a price: what one verified task cost at the provider's public list price β€” CPAT β€” plus how long it took (TTAT).
3 Β· You post a job and a quality bar. The exchange clears at the cheapest model over the bar. That clearing price is the market price of a done task β€” and it moves as models change.
Why only a middleman can run it: pricing a "done task" needs independent verification across every vendor. A vendor can't grade itself; a token reseller doesn't verify. The interface people work through (LLManager) is where that measurement comes from and where the routing gets applied.
πŸ“– Order book β€” models that clear the bar, by price
#ModelGrade (this task)CPATMonthly at your volume
The clearing price is the lowest CPAT among models over your bar. Only we can run this market β€” because pricing a "done task" needs verified measurement, which is exactly the grade. A production venue locks that price for prepaid capacity (futures) and takes the routing spread as margin.

πŸ§ͺ The test battery β€” what actually moves a grade

No opinions, no vibes. Every model runs the same programmatically-checkable tasks. Light tests run constantly so the grade stays fresh; heavy tests run less often but weigh more. Hundreds of these micro-checks roll up into the category scores, and the categories into the single grade β€” that's how small, checkable tasks extrapolate to a macro verdict you can trust.
⚑ LIGHT β€” fast, high-frequency, keep the grade current
πŸ‹ HEAVY β€” complex, weighted, the real separators
Live pass-rates update as tests run above. A test is scored only when its output is machine-verifiable β€” code that must pass hidden asserts, JSON whose totals must add up, a retrieval whose cited fact must be exact. That's what separates a Meerada grade from a leaderboard vote.

πŸŽ›οΈ Meet LLManager β€” it manages every model for you

Not a translator β€” a manager for the communication and tasks between you and your models. Tell it what you want in plain words; it shapes the lean instruction, sends it, checks the answer held up, and keeps the thread. Same result, far fewer tokens β€” on any model you use. See the full product β†’
STEP 1 Β· CONNECT
Your keys, your models
Add your provider keys once. LLManager runs locally and drives Claude, GPT, Gemini, DeepSeek on your own account β€” nothing routed through us.
STEP 2 Β· IT READS YOUR WORKLOAD
Finds where spend leaks
It sees how you actually call models and flags the expensive mistakes: re-sent context that should be cached, flagship models doing work a cheap one clears, runaway reasoning, no output ceiling.
STEP 3 Β· IT RUNS THEM RIGHT
Same result, fraction of the bill
Caches stable context, routes each task to the cheapest model that still passes your quality bar, caps reasoning and output β€” verified, not guessed.
Analyze your workload illustrative Β· LLManager measures your real numbers
Meet LLManager β†’ Free core on GitHub Engine is Apache-2.0 & runs local Β· the cockpit is the premium product.

πŸŽ›οΈ Session Console β€” one cockpit for every model

LLManager is the go-to launcher for any model. Say "open Claude and draft the release notes" β€” it opens Claude, shapes your ask into a lean instruction, and runs it. Fire several at once, across different models, and manage them all from one place instead of juggling tabs. Full product β†’
preview Β· SIMULATION
How it works: LLManager holds your keys locally and drives each provider's API (or app) for you. One plain request β†’ the router picks the cheapest model that clears the quality bar, shapes your words into an efficient instruction, and streams the result back. Parallel sessions share one context so you don't re-paste β€” and one bill's worth of tokens does the work of many.

⬇️ Get LLManager β€” any device

DESKTOP Β· mac / win / linux
CLI + tray app
pip install meerada then meerada up. Runs locally, holds your keys, opens as a menu-bar cockpit. Your prompts never leave your machine.
PHONE Β· iOS / Android
Share-sheet + PWA
Install the web app, or hit Share β†’ Meerada from any app. Speak or paste an intent; it routes to your model and returns the lean answer.
EVERYWHERE Β· bring-your-keys
Your models, your bill
Claude, GPT, Gemini, DeepSeek, local Ollama β€” add a key once. LLManager sits in front of all of them as one interface.
Get LLManager β†’ Free core on GitHub

🀝 The Handshake β€” plan a migration

Pick the model you run today and a candidate. We estimate the savings and quality delta β€” and guarantee, via output equivalence, that you don't lose quality silently.
FROM (today)
β†’
TO (candidate)
Run the real Handshake β†’ Replays your golden set on the candidate, returns a measured per-cluster gap report. Content never leaves your machine.

πŸ“Š Why our grade beats the leaderboards

measures…LMArenaArtificial Analysis$/tokenMeerada
the real cost of a done taskβœ•βœ•βœ•βœ“
counts retries & reasoning burnβœ•βœ•βœ•βœ“
verified success, not preferenceβœ•~βœ•βœ“
updates within the hourweeks~staticβœ“
on YOUR workloadβœ•βœ•βœ•βœ“
A model can win on $/token yet lose on cost per done task β€” cheap tokens don't help if it needs three tries. That gap is exactly what we measure.