FinOps for the AI Era • Nous × NVIDIA × Stripe

PRICE EVERY
TOKEN BEFORE
IT BURNS

A trading floor for your company's compute. Every query is priced, deduped, ranked by value-per-token, and settled against a hard budget — before it runs.

Install on your Hermes — one line, no fork
Install Run a tick
curl -fsSL https://bursar-hermes.com/install.sh | bash
Drops a "Trading Floor" tab into any stock Hermes · prebuilt, no build step
Classical engraving — the bursar weighing value against cost
Bursar — Trading Floor · clearing the market
ORDER FEED — value-per-token, highest first LIVE
SERVICED Reconcile Q2 revenue against the GL export claude-opus-4-8 vpt 8.4
ROUTED ↓ Tag these 400 support tickets by intent gpt-4.1-nano vpt 22.1
DEDUP What's our refund policy for EU customers? cache · $0.00 served free
SERVICED Draft the board memo on the pricing change claude-opus-4-8 vpt 6.9
ROUTED ↓ Summarize this 9k-token meeting transcript claude-haiku-4-5 vpt 14.7
REJECTED Re-run the same vibe-check for the 6th time worth < fee priced out
$41.13SETTLED THIS TICK
$281.07NAÏVE — BEST MODEL FOR ALL
85%SPEND CUT · RECONCILES TO THE CENT
Bursar Trading Floor · settled through StripeHard budget $42.00 · gate held pre-execution
The problem, in three numbers
Enterprise AI budgets

$1.2M → $7M

Average annual spend, 2024 → 2026. Token consumption is up 13× since Jan 2025 — and the alerts still fire after the money's gone.

Near-duplicate queries

~31%

Of production queries are near-duplicates — frontier prices paid to answer the same question five different ways. Research puts 50–90% of inference spend as waste.

What Bursar recovers

85%

Spend cut on a 400-query tick — $41.13 settled versus $281.07 naïve — every dollar reconciling bill = ledger = Stripe, to the cent.

The LoopPre-execution
#1 Price

QUOTED
BEFORE IT
RUNS

Every query is costed at real, sourced market rates — pulled live from Hermes's own pricing snapshot. 28 models, a 528× price spread, no fantasy numbers.

#2 Dedup

WASTE
STARVES
ITSELF

Ask it again and Bursar reuses the earlier work — handing the agent the prior answer as context for one cheap call instead of redoing the whole lookup. It re-verifies the files that answer touched by hash at zero tokens, so reuse stays current and never stale. The 31% never gets paid for twice.

#3 Rank

VALUE PER
TOKEN
FIRST

The dispatcher fills the budget highest value-per-token first — a knapsack market, like a trading desk allocating scarce capital. Best queries clear first.

#4 Gate

A HARD
BUDGET, PRE-
EXECUTION

Per-team tranches with caps that hold before the call is made. A tranche can never be overspent. The overrun doesn't happen — it can't.

#5 Route

CHEAPEST
MODEL THAT
CLEARS

Trivia goes to nano; high-stakes goes to the frontier. Bursar routes to the cheapest model that meets the value tier's bar — killing the "best model for everything" tax.

#6 Settle

A REAL
SETTLEMENT
LEDGER

Every serviced query meters a Stripe trading fee — the metering rail and the throttle. A query whose worth can't clear its cost plus fee is priced out.

Live integration · zero core fork

IT GOVERNS YOUR
REAL TRAFFIC

Bursar isn't just a simulator. A native Hermes agent plugin sits in the LLM execution path of your real chats, subagents, and tool calls — observing, pricing, deduping, and down-routing live. Arm it with a single shield toggle.

BURSAR

Nemotron • Stripe • Hermes • Nous

THE EXCHANGE

One job per rail. NVIDIA's Nemotron values and routes on-prem — your prompts never leave your walls. Stripe settles the ledger. Hermes orchestrates. Bursar is the market between them.

View on GitHub →
BURSAREXCHANGE
A figure weighing a glowing sphere of value
Bursar · the internal compute exchange
B Built on Hermes
Nous · NVIDIA · Stripe