Lab status: workforce rebuilding on measured foundations — qualification program runningread the audit →
KEEN LABS

Scientific AI research lab · Landskrona, Sweden

Every AI company says theirs is best. We measure.

The precision instrument of AI: a 335-model library benchmarked on real tasks, routed on evidence, with a visible receipt under every answer. Chat, Preview and Code on one shared memory. One seat, €49 — and five pilot seats open right now.

Five seats. No card, no spam, one email when your seat opens.

ROUTER · TRACE REPLAYSAMPLE
classify
route
·
llama-3.1-8bdeepseek-flashkimi-k2.6gemini-flash

The panel above replays representative routing decisions — including the tool-call and fallback cases — and is not a live feed or a captured incident log. Standing rule: we publish our findings and method, never our current routing config. The findings are the pitch; the config is the moat.

01 · THE LIBRARY

335 models. Counted, not rounded.

Every Claude, every GPT, DeepSeek, Kimi, Qwen, Grok, Gemini, Llama — held to the same benchmarks, on the same real tasks, every week.

02 · THE WEAVE

Fusion, not favouritism.

The woven core is the router: it can chain a specialist, a verifier and a synthesiser into one answer — and no lab, including the famous ones, gets traffic it hasn’t earned.

03 · THE RAY

One green line: the answer, priced.

The output ray is a receipt — model, milliseconds, sparks. Precision you can audit, on every single answer.

335
models measured — counted, not rounded
122ms
fastest routed answer we have logged
0.11sp
what an everyday question costs — not frontier prices
100%
of answers carry a visible receipt

Act II · the origin · July 2026

We caught our own AI lying.

For three months our autonomous workforce reported success. Then we checked what the verifier was actually doing.

  1. 230completed runsThree months of them. Every one reported as a success.
  2. 75/100the score, every timeIncluding on the runs that had failed. The verifier could not say no.
  3. 150never verified at allNot scored badly. Not scored. The gate was open the whole time.
The row is the proof. No claim without a record. No score without a verifier that can say no.

Most companies would bury that. We paused the workforce ourselves, root-caused it in public, and rebuilt the company on that one rule. Read the full audit →

Act III · what the rule built

One library, 1,014× apart. The router's job: never pay frontier for a greeting — never send a hard question to a budget model.

And it shows its working. Which model answered, how long it took, what it cost — not a settings page you could go and find, printed under the answer, every time. Two of them, out of the same log window:

PRODUCTION RECEIPT · QUICK QUESTIONREPLAYED FROM THE LOG, NOT LIVE

The cheapest kind of answer in the window

model
deepseek-v4-flash
latency
998ms
class
Quick question
charged
0 sp

Charged nothing at all. A quick question does not need a model that can reason for thirty seconds, and it did not get one.

PRODUCTION RECEIPT · REGULATED GUIDANCEREPLAYED FROM THE LOG, NOT LIVE

The dearest answer in the window

model
gpt-4o-mini
latency
3.4s
class
Regulated guidance
charged
0.8 sp

The most expensive single answer in the whole snapshot — 0.8 sp, for guidance that has to hold up against a regulated source.

Same subscription, same week — the router priced each question on its merits.

A replay of two real rows from the production cost log (2026-08-01 to 2026-08-18) — both appear in the measured table further down this page, one of 48 in its class and one of 2 in its. The cards are a rendering, not a live feed and not a captured session.

The rest of the range · measured, not charged

Cheap, working, frontier — and what each one is for.

Our traffic so far is cheap — that is the router doing its job. What it reaches for when a question earns it is measured below: all 54 models on one task set, one card per price tier, cheapest first. These are the things the router is picking between.

Budget tier$0.0001 – $0.001 / task

Greetings, lookups, one-line answers.

gpt-5.6-lunaOpenAI

$ / task
$0.000886
latency p50
2033ms
pass rate
91.9%
tasks scored
124

Row the price ceiling of this tier — the dearest of 9 models measured between $0.0001 – $0.001 per task, not the best of them.

Tier 1,009 scored tasks, 80.6% median pass, 1296ms median p50.

MEASURED · QUALIFICATION SWEEP · 2026-08-17

Working tier$0.001 – $0.01 / task

Real tasks — drafting, review, structured output.

claude-sonnet-4-6Anthropic

$ / task
$0.00987
latency p50
4702ms
pass rate
90.5%
tasks scored
21

Row the price ceiling of this tier — the dearest of 29 models measured between $0.001 – $0.01 per task, not the best of them.

Tier 1,216 scored tasks, 85.7% median pass, 3492ms median p50.

MEASURED · QUALIFICATION SWEEP · 2026-08-17

Frontier tier$0.01 – $0.1 / task

Reserved for the moments that earn it.

claude-fable-5Anthropic

$ / task
$0.0436
latency p50
4623ms
pass rate
72.7%
tasks scored
11

Row the price ceiling of this tier — the dearest of 13 models measured between $0.01 – $0.1 per task, not the best of them.

Tier 172 scored tasks, 90.9% median pass, 3881ms median p50.

MEASURED · QUALIFICATION SWEEP · 2026-08-17

These are benchmark rows, not receipts. Nobody was charged for them and nobody asked for them: they are graded tasks out of the qualification sweep — one task set, 11 domains, every model under identical conditions — so they carry a cost in dollars per task and no Sparks figure at all. The two receipts above are the other series: real answers, real people, real charges. The line at the top of each card is what we route that band for; the figures under it are what the band measured.

Proof · quality gate

Loading reviewer outcomes from ai.keenlabs.pro…

Method

Measure. Route. Re-measure.

01 · MEASURE

Real tasks, scored

Every model benchmarked on user-shaped work — build, repair, analyse — graded by verifiers that can say no.

02 · ROUTE

Evidence picks the model

Quality per dollar decides. A cheap model that wins gets the traffic. A famous one that doesn’t, doesn’t.

03 · RE-MEASURE

Every answer feeds back

Receipts, fallbacks, comparisons — your usage is our research. The system gets measurably better the more it’s used.

Skill cards are dated, versioned, and regenerated whenever a model updates. Status is reported with row counts, not adjectives: zero rows means not started, however good the design document is.

The living library

335 models, ranked by evidence, re-ranked as it changes.

Every Claude, every GPT, DeepSeek, Kimi, Qwen, Grok, Gemini, Llama — plus live tools no chatbot carries: market data, web search, weather, cited sources. The router serves whatever wins today — promoted and demoted by data, not by brand.

The current active rotation — 12 of 335 measured. Of those, 54 have been run against the same task set under identical conditions — that benchmark is the section below, and the gap between 54 and 335 is work still to do.

The spread · measured

One task set, 11 domains, 1,014× apart in price.

52 priced models, every one of them run against the same tasks, with cost measured from the tokens each one spent. From the floor of the library to the ceiling is a factor of 1,014. Choosing between them, per question, is the whole job — and what the router chose for real questions is the section below.

COST PER 1,000 TASKS · 52 PRICED MODELS · ONE TASK SETMEASURED

One tick per priced model, cost per task. Logarithmic, because the library spans 4 decades and a linear axis would draw 51 of these as hairlines beside the dearest one. 2 further models measured at zero — a free tier, not a missing figure — and cannot sit on a log axis: gemma-4-26b-a4b-it and gemma-4-31b-it, which scored 31.6% and 42.4%.

  • Cheapest priced modelGroq · llama-3.1-8b-instant
    $0.043

    64.0% pass over 114 scored tasks · $0.000043 per task · 10 of 11 domains. The floor of the measured library.

  • Cheapest above 90% passGoogle · gemini-3.1-flash-lite
    $0.71

    91.1% pass over 124 scored tasks · $0.000714 per task. The cheapest row in the sweep that cleared the threshold — and it cleared it on the full task set, not on the spine.

  • Dearest model measuredAnthropic · claude-fable-5
    $43.59

    72.7% pass over 11 scored tasks · $0.0436 per task. Eleven tasks is the core spine only. Read that pass rate with the depth, in both directions.

$0.71 against $43.59 per thousand tasks — 61× — on work graded the same way. The cheapest model that cleared 90% did it over 124 scored tasks; the dearest model in the library ran 11 and passed 72.7% of them. End to end the priced library spans 1,014×. Compare only the 15 models that all ran the full suite (95124 tasks each), where the means are taken over comparable work, and it is still 77× — with the dearest of those, qwen/qwen3.6-27b, holding the lowest pass rate in the entire sweep at 15.7%. Price does not predict the score in either direction. That is the routing problem, measured.

What this is, and is not. Measured 2026-08-17, suite qualification-v1, 54 models on one task set of 11 domains — not list prices. Costs are computed from measured token counts at published provider rates: arithmetic on a measurement, not an invoice, and they exclude negotiated and free-tier pricing. Coverage is not quite even either. 50 of 54 models returned a scoreable answer in every domain; the other 4 did not, and one of them is the cheapest row above — so their domain counts are printed on the rows rather than averaged away. Depth is uneven. Cost per task is a mean over that model's run, and the runs range from the 11-task core spine to the full set, so this is cost per task within one suite rather than a paired task-by-task comparison — the export publishes cost per model, not cost per task per model. One run is a snapshot.

Full per-model table, failure modes and the stated limits: /research#results

What the router chose · real traffic

And here is what it actually picked, for questions real people asked.

145 answers out of the production cost log, 2026-08-01 to 2026-08-18. Every row of it is cheap, and that is the finding rather than an embarrassment: nothing anyone asked in this window needed the top of the range, so the router did not buy it. This is a picture of our traffic, not of our library. The library is the section above.

Measured · 2026-08-01 to 2026-08-18LatencyCharged
llama-3.3-70b-versatile · Code (n=20)376ms0 sp
deepseek-v4-flash · Everyday question (n=72)936ms0.01 sp
deepseek-v4-flash · Quick question (n=48)998ms0 sp
llama-3.3-70b-versatile · Research (n=1)1.5s0.5 sp
llama-3.3-70b-versatile · Long analysis (n=2)2.6s0.2 sp
gpt-4o-mini · Regulated guidance (n=2)3.4s0.8 sp

One median-latency row per task class with n beside it, and the Sparks column is what a customer was actually charged — these are receipts, not benchmark rows, and the two series are never mixed. A snapshot, not advice: by the time you read this, the data may have moved it. That’s the feature. Standing rule: we publish findings and method — never the current routing config. The findings are the pitch; the config is the moat.

Proof of work

Built by the same instruments we sell.

LIVE

Keenoble

The flagship — smart fusion routing across Chat, Preview and Code at ai.keenlabs.pro.

Open surface ↗

LIVE

Pensionskollen

Swedish pension and payroll guidance behind regulated source guardrails. Same engine as LöntagarPension, two doors.

Open surface ↗

LIVE

MC Garage Landskrona

A real-world business running the workforce pattern — bookings, keys, revenue. Built with Areka Capital AB.

Open surface ↗

QUEUED

WSP Screener

Market screening intelligence, queued for workforce-backed research loops.

Keen OS Command · by invitation

For companies that need more than a seat: custom recurring agent systems on the same foundation — source watches, eval loops, PR drafting, verification gates, human approval on everything customer-impacting. Deployed in engagements from €2,500, not sold on a pricing page. It wasn’t built to demo well; it was built to run.

Tell us what you’re building →

Act IV · one project · three rooms

Chat, Preview and Code — born connected on one memory.

ROOM 01

Chat

Say it once.

Every prompt routed to the model the evidence ranks best — or to a live tool when a model would only guess.

  • Speed for small things, depth for hard ones
  • A receipt under every answer
  • Compare models side by side
ROOM 02

Preview

It already knows.

Idea → live preview → deployed, without re-explaining. The preview inherits the goal, the constraints, the half-thoughts.

  • Context flows in from Chat
  • Publish public or private
  • No deploy infrastructure to think about
ROOM 03

Code

The why is attached.

Surgical, repo-connected control with the full project context attached to every file you open.

  • The 90 % AI got right, plus your 10 %
  • GitHub connect
  • Ask and patch against shared memory

Say it once in Chat. Preview already knows when you arrive. Open Code — the source is there with the why attached. Visible memory, editable in one tap.

PROJECT MEMORY · SAMPLE MAP — nordic-saas/
CHATPREVIEWCODE
index.md · mapfoundation/stackdecisions/dark-themepreferences/brandfiles/pricingartifacts/deploy
You say in Chat: “Brand colour is emerald, keep the tone dry and Nordic.”
wrote preferences/brand · linked from index.md

Projects are folders of linked memory nodes — an index map, decisions, preferences, files. Each surface follows a path to exactly the node it needs. Visible, editable, yours. Sample project, drawn to show the mechanism.

One memory layer across all three — structurally impossible for five separate subscriptions.

Mission

The precision instrument of AI — open to ordinary people.

The highest intelligence should not pool inside a handful of private companies. We buy every model, measure them all honestly, and route your work to whatever the evidence says is best — because no lab selling its own model can afford to tell you when a rival’s is better. We can. That’s not a claim about anyone’s engineering. It’s business-model logic.

Neutrality (no lab favoured), receipts (no hidden costs), open findings (method public, config private), and the €49 price point are not marketing choices — they are mission constraints binding the engineering. Read the mission in full.

What we don’t claim

No revenue yet.Zero paying customers. We prove the model works end-to-end before charging anyone.
No independent benchmarks.Every measurement is ours — method published, receipts dated. Verification welcome: hello@keenlabs.pro.
We never assert super-intelligence.It’s a word the measurements may one day earn. What we assert, and prove with rows, is the democratisation itself.
No silent output caps.Every model runs at its full documented ceiling. Ask your current platform what its hidden limits are.
The Truth Auditor only sees what it sees.A verifier can only catch what it checks — that limitation is why the audit exists.
No uncited figures.A number with no source and date is a bug — tell us and we fix it in public.

The math

Cancel the stack.

Replicating what one Keenoble seat gives you costs $105 a month across 5 logins whose tools have never met each other. Then add pay-as-you-go API keys for the models those subscriptions don’t include — DeepSeek, Kimi, Qwen, Grok — and live data feeds on top. You’re past $150/month and you still have zero shared memory, no routing, and five tabs of copy-paste.

Without Keen

$105/mo + API keys

ChatGPT Plus$20
Claude Pro$20
Cursor Pro$20
Lovable Pro$25
Perplexity Pro$20
5 logins, zero shared memory$105+

With Keenoble · Premium

49/mo · 1,000 Sparks

8,708 everyday answers at the 0.1148 sp we actually charged, or 1,250 at the dearest answer we have logged (0.8 sp). Measured, 2026-08-01 to 2026-08-18 — not a forecast of your month.

  • Chat — smart-routed across the measured 335-model library
  • Preview — idea to deployed, context inherited
  • Code — repo-connected, shared memory
  • Live tools — market data, web, weather, sources cited
  • Receipts on every answer · cancel anytime
Join the pilot programme →

One currency, one receipt: you pay for work you actually received, and a run that fails its gate doesn’t bill. Keenoble is designed to consolidate AI tools for thinking, writing, coding and building. Image, video and music generation are already catalogued in the lab and will be released as credit add-ons on top of Premium — not claimed as shipped until they are. Competitor prices from public pricing pages, 2026-05-21.

Pilot · five seats

Five seats. That’s the whole programme.

One month of Premium, free, with a direct line to the founder — you tell us what breaks and what’s brilliant, we ship fixes while you watch the changelog move. We picked five because we actually mean the direct line.

1
2
3
4
5

All five open. Updated by hand — there is no live counter here.

Five seats. No card, no spam, one email when your seat opens.

Support

Quick answers, honest ones.

What are Sparks, exactly?

The unit you pay with. Every answer shows its price in Sparks on the receipt — a greeting is a fraction of one, a deep reasoning task a few. Premium includes 1,000 Sparks a month and you can top up if you run out. One currency, one receipt: you pay for work you actually received.

Which model answers my question?

Whichever the evidence ranks best for that task today — and the receipt tells you which one it was, every time. You can also pick manually from the library. Read the findings →

Can I cancel anytime?

Yes — Premium is a monthly plan with no lock-in, and the cancel button works exactly like the buy button. No retention flows, no dark patterns.

What happens when a model fails mid-answer?

The router falls back and the receipt shows the correction in public — you pay for what answered, not what we hoped would.

What do you publish, and what do you keep back?

Method public, config private. We publish our findings, our method and our mission — including the audit that caught our own system reporting success it hadn’t earned. We don’t publish the current routing configuration. The findings are the pitch; the config is the moat.

More questions →

The builder

Keen Andersson Lang

Founder, Keen Labs (2026). Landskrona, Sweden. Payroll Business Partner student (Stockholm School of Business), incoming at ASSA ABLOY autumn 2026 — building Keen Labs alongside. Regulated processes are where wrong answers cost real money; that’s where the lab’s verification obsession comes from.

Built this lab solo, with an AI workforce as the team — and when that workforce lied, audited it in public. If you’re going to trust an AI company, trust the one that showed you its worst finding unprompted.

Keen Andersson