Skip to content

Usage & billing

LunaRoute coding plans are flat-rate for inference. Your lanes are the limit and the price is the price: there is no token cap, no inference usage meter to watch, and no overage billing for model calls. Heavy inference use doesn’t trigger usage-based blocking or extra charges. When the fleet is busy, it runs at a lower fair-share scheduling priority and may wait longer to start.

A few agent tools are metered separately. web_search over MCP includes a sponsored monthly allowance, and platform-key searches beyond it accrue to the usage ledger and are charged on the next cycle.

Current plan names and prices are on www.lunaroute.com. A one-lane plan for a personal agent is available on the personal-agents page.

Every request is recorded in the dashboard with its billing metadata — timestamp, model, token counts, and cost — so you can see where spend goes. On flat-rate inference this is for visibility, not a bill; separately metered MCP usage (above) is the exception.

Request and response bodies are not stored — see Zero data retention. Only billing metadata is kept.

Upgrades take effect immediately and prorate; downgrades start the next cycle. Your effective lane caps update with the plan (and any organization override).

See the dashboard for your plan, ledger, and invoices.