Skip to main content
The monetization layer for usage-based products

Usage, priced and enforced at runtime.

Limitr runs inside your app to price, gate, and meter every unit of usage in real time, so margin holds across every plan, contract, and custom deal.

The hard part

Every price change is a release.

Model costs, usage mix, and contract terms shift constantly, so pricing needs constant adjusting, like a sail in shifting wind. Spread across your front end, back end, and billing, every adjustment means changing them all together.

The fix

One policy shapes every unit of usage.

A versioned document — JSON, YAML, TOML, or Stof — declares plans, entitlements, prices, limits, and rules once, managed from Limitr Cloud. The engine runs it inside your process as WebAssembly, so every service reads the same policy, and a change goes live everywhere without a release.

The difference

Margin is engineered across calls, not billed per call.

One outcome is dozens of calls: models, tools, retries. A billing API sees each one alone. Limitr holds the running state in your process, so it prices and limits the whole run as it happens, against live cost and margin.

  1. Too looseleaks margin
  2. Too tightstalls adoption
  3. Trimmedmargin holds
Built for custom contracts

Every contract, enforced as written.

Each customer’s terms live in the same policy and apply on their next call. Sales or finance can set them up, no new plan, no engineering ticket.

  1. Acme20% off outcomes through Q4
  2. Globex$2,400 monthly spend cap, hard stop
  3. Initech$50k annual commit, drawn down per call
The Regis Company
Customer story

The Regis Company chose Limitr to meter and enforce real-time AI avatar usage ahead of a Fortune 500 rollout covering more than 100,000 users.

Fortune 500
Enterprise rollout on a fixed-price contract
100,000+
Users on a single contract
3 modalities
Chat, text-to-speech, and real-time avatars
Read the customer story
Where you plug in

The runtime between the contract and the invoice.

Terms are agreed in your CRM or at signup, and revenue is booked in your ERP. Limitr, inside your app and managed in Limitr Cloud, is the piece in between that actually runs them.

  1. 01
    Terms agreedSigned in your CRM, or picked at signup.
  2. 02
    Terms become policySet once in Limitr Cloud, pushed live.
  3. 03
    Enforced in your appThe Limitr engine runs every call locally.
  4. 04
    Recorded and invoicedUsage events land in the Cloud ledger.
  5. 05
    CollectedStripe out of the box, or your own rails.
  6. 06
    Revenue bookedSynced to your ERP or warehouse.
In your code

One document. A few lines to read it.

  1. DefineWrite the policyPlans, prices, and limits in one file.
  2. LoadRun it in-processOne call starts the engine.
  3. AskCall allow()Gate, meter, and price usage in one step.
  4. TrimChange it liveEdits publish everywhere. No deploy.
Policy
plan.starter · v1
EntitlementIncludedPriceMode
Seats3$19.99hard
AI Chat1,000 / hr$0.00004 / tokensoft

Published. Every connected engine reads it on the next call.

Integration@formata/limitr
import { Limitr } from '@formata/limitr';

const policy = await Limitr.cloud({ token });

await policy.createCustomer('acme_corp', 'starter');

if (await policy.allow('acme_corp', 'seats', 2)) {
  console.log('2 seats — allowed');
}
if (await policy.allow('acme_corp', 'ai-chat', 10000)) {
  console.log('10,000 tokens — allowed, 9k billed as overage');
}
Open source

The same engine, free and self-hosted. Commit the policy to your own repo and run it anywhere WebAssembly runs.

View on GitHub
Limitr Cloud

Adds the hosted policy editor, version history with one-click rollback, and billing, ledgers, and margin analytics. Moving over is a one-line change.

Cloud quickstart
What are you trying to do?

Every feature can run its own pricing model.

UsageSeatsHybridCreditsTrialsRolloversOutcomesEnterprise deals

Attribute every AI agent’s usage, automatically.

Per vendor, feature, contract, and customer.

Enforce and bill every API call, live.

Gated, metered, and priced in one step.

Cap what agents calling your MCP server can spend.

Shared credit pools with hard ceilings.

Cut a custom rate for one enterprise customer.

An override with an expiry. No new plan.

Reprice a single feature, not the whole model.

One entitlement changes. Nothing else moves.

Get alerted when a usage threshold is crossed.

Trial limits, upsell triggers, runaway spend.

Why Limitr

A monetization layer, not just a billing API.

µsvs. 50–100+ ms per remote call

In your process

Decisions run locally, even offline. Usage syncs in the background and never blocks the hot path.

1 docversioned · diffable · committable

A document you own

Review pricing like code. Cloud adds history, instant publish, and one-click rollback.

cost + priceon every credit

Margin, live

Profitability per customer, feature, and vendor, continuously. Not once a month at close.

What you get
No-deploy price changesChange a price or tier live. Nothing to ship.
Custom deals & overridesOne-off rates, discounts, and rules per customer.
Any customer typeOrgs, users, agents, and workflows, governed the same way.
Live alertingNotified the moment a limit or anomaly hits.
Credit top-upsCustomers buy more before they run out.
Extend the runtimeCustom logic and units when defaults aren’t enough.
CJ and Amelia, co-founders of Limitr
“We wanted a config document for all things usage and pricing — something we could just version and commit. That’s why we built Limitr. That philosophy doesn’t exist in the other products, even when their marketing sounds the same.”
CJ · Co-Founder & CEO
Why teams choose Limitr

Infrastructure that holds up in real conditions.

Five reasons teams pick Limitr once pricing meets real customers, real contracts, and real-world conditions.

  1. Harbor
    Experienced team, not just a platformWe work with your team on pricing strategy and unit economics.
  2. Wind shift
    Flexible monetization per featureGive any feature its own strategy, when it needs one.
  3. Beacon
    Usage intelligence, built inLive alerts, forecasting, and AI guidance on pricing changes.
  4. Squall
    Durable in real-world conditionsHolds up through usage bursts, with offline behavior you configure.
  5. Open water
    You never outgrow itOverride built-in rules and add your own logic; Limitr Cloud tracks the rest.

Common questions

Start here

A billing platform records what happened and charges for it after the fact — invoices, payments, subscriptions, revenue recognition. Limitr decides what's allowed before the action executes, and meters it in the same operation.

If enforcement lives downstream of consumption, the cost is already incurred by the time anything reacts. Limitr runs at the moment of consumption, in-process, so a limit holds before you pay a vendor for the tokens rather than after.

Keep Stripe, Maxio, QuickBooks, or your homegrown invoicing for collecting money. Limitr is the monetization layer above it: one policy document that decides what each customer gets, what it costs you, and what you charge, enforced live. We integrate with the billing platform and workflow of choice.

No. Limiting access and charging for usage are separate decisions your policy can make independently. Most teams start by just watching usage — you can't control what you can't see — then add limits, and monetize only the parts that make sense.

The usage you charge for doesn't have to be the same usage you control — most teams layer a few entitlements together to get the pricing and access model they actually want (e.g. control tokens, charge only for successful runs).

Yes — usage-based, seat-based, hybrid, credits, trials, rollovers, custom enterprise deals. One policy format expresses all of it, so if you can describe the model, you can ship it without rebuilding your billing logic for the next one.

That's the point of a document instead of hardcoded logic: pricing models keep changing, the format doesn't have to.

No. AI is the loudest version of the problem, not the whole problem. Limitr meters any unit you can define: AI tokens by model, API calls, seats, vendor connections, SMS messages, avatar minutes, GPU seconds, storage, agent runs, outcomes, or a composite credit of your own design.

In practice our customers split roughly evenly. Some are metering multi-model AI pipelines. Others are legacy SaaS companies moving off flat subscriptions — per-seat billing that keeps drifting out of sync, text-message overages reconciled by hand, or a shift from a complicated spend-based calculator to a clean per-unit price. Same infrastructure problem, no AI involved.

All three. Plenty of tools will tell you what your customers consumed after the fact. Limitr enforces the limit at the moment of consumption, meters usage in that same atomic operation, and feeds the result straight into invoicing.

Limits run in three modes, switchable without a code change: hard blocks the action at the cap, soft allows the overage and optionally bills on it, and observe tracks everything while enforcing nothing. Most teams start in observe, learn what their usage actually looks like, then turn on enforcement once the data supports the decision.

For engineering

No. The policy runs in-process as WebAssembly, colocated with your application. Every allow(), check(), and increment() executes locally in microseconds, with no network call on the hot path — compared with the 50 to 100+ milliseconds typical of enforcement systems that make a remote API call per decision.

Usage events sync to Limitr Cloud asynchronously over a background WebSocket, and policy updates push down the same way. The hot path stays fast; your data stays current.

The integration surface is deliberately small. You call ensureCustomer at login or first touch, then policy.allow at the point of consumption with an entitlement name and an amount. That's the core of it.

Engineering wires up the entitlement by name and passes in usage quantities (optionally gating access, too). Everything behind that name — limits, prices, tiers, per-customer overrides — is configuration that changes without touching your code.

The SDKs are each a thin wrapper over the same core WASM module, so additional languages take days rather than quarters to support. TypeScript/JS is most supported today, but tell us what you're running and we'll confirm timing before you commit to anything.

Dynamic model switching is a core part of AI usage control, and we are very good at that. We are also very good at putting prices to usage, so you can track and control margin-to-serve in the ways that matter, not just a general aggregate cost-to-serve which leaves a lot of blind spots.

Limitr does not, however, make LLM requests on your behalf, switch models, or repackage specific vendor tokens. We don't sit in the request path routing calls — we provide a control plane around them, no matter which vendor or router they end up going through. Limitr works with any AI vendor, including alongside a router.

If you're looking for spend management and margin control, Limitr will suffice on its own. If you're also looking for a unified endpoint, model routing, & other AI-specific guardrails, a router may make sense running alongside it.

On disconnect, the runtime will try reconnecting automatically and gives you the option to continue operating normally offline or block all usage until back online (the default to preserve state). If your app is offline, you typically have larger issues to contend with. The default is configurable in the SDK, and enforcement never stops working.

The open-source engine also runs fully offline against a local policy file, enabling you serve users in low or no-connectivity regions, or have a separate offline/backup usage policy for failure and testing modes of operation.

For bad data: meters move in both directions. If a runaway process inflates a counter, decrement it through the SDK or set the value directly in the dashboard. Counting happens locally and serializes to the server, where operations are atomic, so you won't double-count.

Same core. The open-source enforcement engine runs in-process, is fully self-hostable, and is free permanently. Cloud adds the managed policy dashboard, versioning and instant rollback, per-customer analytics and margin data, usage-based invoicing, live alerting, and the Stripe/billing integrations.

The enforcement API is identical either way, so moving from local to Cloud is a one-line change. Prototype against the open source engine without talking to us first — several teams have.

Yes, our open-source runtime built on Stof contains validation rules, and can literally validate itself so that you never have an invalid policy.

Inside the npm.js package, Limitr.new(..) actually calls validate by default and will throw an error if the provided policy does not fit the spec, providing a message for where the error resides.

If you're using Limitr Cloud, the policy is fully managed and will always be validated when changes are made and before it gets used within your app or service.

Less than you'd expect. Every enforcement decision runs in-process, so your requests, prompts, and responses never pass through Limitr. With Limitr Cloud, what syncs is usage: which customer consumed which entitlement, how much, and when, plus the customer IDs and any metadata you choose to attach.

Run the open-source engine on its own and nothing leaves at all: the policy is a local file and enforcement works fully offline. For Cloud customers with GDPR or similar obligations, we can put a Data Processing Addendum in place, and we're glad to walk your security team through the architecture.

For finance and product

Yes, and it's the main reason teams buy this. Anyone with dashboard access and permission can change a limit, add a tier, adjust an overage rule, apply a discount, or gate a feature. Changes publish live and propagate to every connected service instantly — no PR, no deploy, no release coordination.

Per-customer and per-segment pricing works through rules rather than proliferating plans. Keep a small number of plans, then layer overrides on top: replace a price outright or apply a percentage in either direction, with an expiry date and an approval note attached. Sales can structure a custom deal without engineering in the loop.

Every change is versioned with full history and reverts instantly.

Honest answer: if your pricing model is static, you're just trying to monetize, and hardcoding it is fine, then you should. The case for Limitr is a function of how often your packaging changes and how expensive it is to extend, maintain, customize, and analyze. Transparently, though, almost nobody's model stays static, and once you dive in, this critical layer of your app gets large, messy, and nuanced, quickly.

It adds up once you need to split billing across products, apply a discount to one specific tool call, burndown credit balances differently across accounts, run different budgets per model in the same pipeline, handle a mid-period upgrade without double-charging, or roll out a tier without a deploy. Together they become a permanent engineering surface that grows with every commercial decision the business makes, often brittle and expensive to maintain.

The estimates we hear from technical buyers cluster around 1-3 engineer-months for a first version, plus indefinite maintenance and additional internal projects for observability. Instead, our contracts are priced for outsized savings on your end, and because the enforcement engine is open source, you aren't betting your pricing infrastructure on our roadmap.

Complexity is the use case. Customer objects are arbitrary: an org, a workspace, a user, an agent, a project, whatever you need to meter. They reference each other, so a shared meter like seats resolves to the right level automatically, and group invoices roll up from individual ones. Credits carry both what you pay a vendor and what you charge, per model and per vendor, so margin resolves per customer, per feature, and per provider without extra instrumentation.

On timing: not too early. Observe mode exists precisely for this — meter everything, enforce nothing, and set your allotments from real data instead of a guess. Implement in stages over time, when and how it makes the most sense for your specific products.

Yes — margin is a first-class output, not something reconstructed after invoices close. Because Limitr already knows what a call costs you and what you charge for it, it can show profitability per customer, per feature, and per vendor continuously, not once a month at close.

That's what the dashboard's margin view is built around: the same policy that enforces limits live also prices your unit economics live.

The open-source engine is free, permanently. Limitr Cloud is contracted annually or on custom terms, and scoped with you on a call rather than from a public price list, so it fits your model instead of making you fit a tier.

If you'd like to evaluate first, start on the open-source engine. Moving to Cloud later is a one-line change.

You keep your pricing. The policy is a plain document (JSON, YAML, TOML, or Stof) that you can export and commit to your own repo, and the engine that runs it is open source. Leave Cloud and you can keep enforcing the same policy on the self-hosted engine, through the same API, so your integration doesn't change.

Configurations and usage records are exportable at any time, and for 30 days after an account closes.

Next step

Turn usage into decisions, not just invoices.