Aurona.ai

Pricing

Start with credits. Scale with commitments.

A developer-friendly path from free testing to AI token credits, production routes, Fusion orchestration, enterprise policy, and GPU-backed capacity.

OpenAI-compatible

Token ledger

GPU route layer

Start

Free

Scale

Credits

Enterprise

Custom

Credit formula

model cost + Fusion orchestration + policy controls + GPU capacity + app margin

transparent model price

credit wallet

budget keys

capacity commitments

Pricing signals

The credit model should reveal exactly what customers pay for.

Aurona pricing should separate raw model cost, route intelligence, Fusion quality, enterprise governance, app margin, and GPU capacity.

Pass-through

model cost

Expose public model input and output pricing as the baseline for transparent credit conversion.

Route value

Fusion fee

Price the orchestration layer that selects panels, judges outputs, applies policy, and returns traces.

Margin control

budget keys

Map usage to apps, customers, teams, environments, and route ceilings before traffic scales.

Compute premium

GPU lanes

Plan serverless endpoints, burst pools, reserved lanes, dedicated endpoints, and next-generation GPU capacity as separate commercial paths.

Credits ledger

balance + coupon

Keep paid credits, starter credits, enterprise credits, coupons, and partner settlement lines distinct.

Credit bundles

planner only

Model starter, promotional, top-up, and committed-credit scenarios without publishing a binding rate card.

Key mode

BYOK or managed

Show when customer provider keys, Aurona-managed supply, or hybrid fallback changes the commercial object.

Planning

simulator

Estimate model, Fusion, fallback, app attribution, broadcast logs, and GPU exposure before commitment.

Cache accounting

read + write

Separate provider-reported cache status from fresh input so repeated prompts and tool schemas remain auditable.

Key control

usage tabs

Expose per-key spend charts, budget progress, log shortcuts, and rotation state inside the workspace.

Credit calculator

Estimate token credits with real controls

250,000

billing modecredit wallet
volume multiplier92%
hard stopcustom budget key
key modeAurona managed
observabilitybroadcast enabled

Estimated monthly credits

6

Builder

model tokens

5

Fusion

2

GPU lane

0

provider vault

0

broadcast logs

0

{
  "plan": "Builder",
  "messages": 250000,
  "model_tokens": "4.50 credits",
  "fusion": "1.75 credits",
  "gpu_lane": "0.00 credits",
  "provider_key_mode": "aurona_managed",
  "observability_broadcast": "0.25 credits",
  "credit_adjustment": "0.00 credits",
  "endpoint_health": "not_selected",
  "volume_multiplier": 0.92,
  "estimated_total": "6 credits"
}

Request reconciliation lab · simulated

Review delivery before a request becomes a ledger line.

Trace provider attempts, usable output, fallback recovery, auxiliary work, and provisional AI Token exposure. This planning view does not create a billing policy or issue credits.

Request outcome

Reconciliation preview

Usable text response

aurona/auto-fast · serverless route

delivered preview

Attempts

1

Usable output

observed

Modeled exposure

0.038 AI Tokens

Provider attempts

  1. Attempt 1serverless

    approved text pool

    usable output

Provisional lines

model input + output

0.032

provisional usage

route evidence

0.006

provisional usage

  1. 01request received
  2. 02route evaluated
  3. 03output checked
  4. 04usage reconciled

Simulation only. This preview does not decide charges, issue credits, change billing policy, allocate capacity, or guarantee a refund.

receipt · reconcile:usable-text:fallback-on:0.038

Usage Settlement Lab · simulated

Translate mixed usage without hiding the native meter.

Model an AI Token reservation across text, image, audio, or video, then reconcile only when complete provider-reported usage is present.

reservation preview

Proposed preview

/v1/credits/settlement-preview?workload=text&tier=shared

Capacity

serverless

Receipt

settlement_00vquer1

Modeled reservation

4.02 AI Tokens

Settlement preview

Pending evidence

input

native token meter

1,200 tokens

cached input

native token meter

300 tokens

reasoning

native token meter

180 tokens

output

native token meter

240 tokens

01

request estimated

02

AI Token reservation modeled

03

provider usage pending

04

wallet unchanged

Simulated translation weights and usage only. This lab does not debit AI Tokens, change a wallet, publish a rate card, or reserve GPU capacity.

Credit Headroom Lab · simulated

Review spend and concurrency before a request group enters a lane.

Join AI Token balance, a workspace reserve, in-flight demand, and modeled GPU-backed route headroom in one admission receipt.

admit preview

Planning receipt

headroom_b4565b26

POST /v1/credits/headroom-preview

Spendable

100 AI Tokens

Request estimate

12 AI Tokens

Lane headroom

4 after preview

AI Token reserve and modeled lane headroom both remain reviewable.

Modeled ceiling: 18 · current in flight: 8 · requested: 6

01

balance evidence reviewed

02

workspace reserve preserved

03

request exposure modeled

04

lane admission previewed

Demonstration values only. This preview does not debit AI Tokens. It does not set a production rate limit, promise throughput, change billing, or allocate GPU capacity.

Plans

Pricing that matches how AI products grow.

Aurona can support early developers, production teams, AI app companies, enterprises, and compute partners without changing the underlying API.

Developer

Free to start

For prototypes, model testing, and early route integration.

  • API keys
  • model catalog
  • Chat playground
  • usage logs
  • starter credits
  • coupon ledger

Builder

Token credits

For teams shipping production traffic with transparent model costs.

  • credit wallet
  • route aliases
  • fallback
  • budgets
  • webhooks
  • endpoint tests

Growth

Committed credits

For apps that need predictable volume, margin control, and support.

  • volume discounts
  • app attribution
  • Fusion presets
  • broadcast logs
  • priority lanes

Enterprise

Custom

For governed usage, procurement, private policy, and reserved capacity.

  • SSO/SAML
  • ZDR
  • audit logs
  • sovereign routes
  • GPU commitments

Workspace billing

Charge by credits, govern by workspace.

A user starts with a personal workspace, then upgrades into a shared company or enterprise workspace when multiple people need the same keys, budgets, logs, data policies, and invoices.

Workspace

Billing basis

What it manages

Commercial motion

Personal workspace

Free sandbox + starter credits

one owner, project API keys, playground, usage logs

upgrade to Company when a team needs shared billing

Company workspace

Pay-as-you-go AI Token credits

members, roles, projects, budgets, invoices, customer attribution

set monthly caps and route-level hard stops

Growth workspace

Committed credits

volume planning, app margins, priority routes, usage reporting

reserve credits for predictable production traffic

Enterprise workspace

Custom contract and invoice

SSO, audit logs, ZDR, data policy, procurement, SLAs

attach private routes or reserved GPU capacity

Compute partner workspace

Capacity settlement

GPU inventory, SLA, DD evidence, facility data, network profile

settle approved supply through capacity agreements

Credit meters

Token credits become the commercial unit.

Different AI workloads can settle through one wallet while preserving model-specific, route-specific, and compute-specific economics.

Workload

Meters

Pricing drivers

Model tokens

input + output tokens

public model price, provider, context, volume

Cached input

cache reads + cache writes

provider-reported cache status, route affinity, reusable context, and miss reason

Starter and promo credits

coupon, trial, or enterprise allocation

expiration, workspace, and invoice treatment

Credit bundle scenarios

starter, top-up, committed, or partner credit plans

non-binding planning view before billing approval

Fusion routes

panel + judge + final tokens

panel size, judge model, trace level

Multimodal apps

media units + tokens

image/audio/video provider and output size

Tool and agent runs

tool calls + tokens + retries

workflow depth and fallback policy

Provider vault

customer-owned provider key usage

key mode, fallback, trace, and route margin

API key controls

per-key usage, budget, and log filters

workspace, route, environment, allowlist, and rotation state

Tier limit previews

request volume, token volume, modality, endpoint mode

account governance and launch review

Broadcast logs

trace fanout events

logs, webhooks, billing exports, analytics sinks

Serverless endpoints

endpoint seconds, warm-cache tests, queue samples

GPU class, region, quota, and runtime mode

Dedicated inference

reserved capacity + usage

GPU class, region, throughput, commitment, endpoint mode

Commercial layers

Make the bill explain the product.

A great pricing page should show exactly what Aurona is monetizing: model access, routing intelligence, Fusion quality, policy governance, and GPU-backed capacity.

Layer

Pricing basis

Value delivered

Public model access

pass-through model cost + route margin

transparent source pricing

Aurona Auto

routing and fallback fee

cost, latency, policy, provider health

Aurona Fusion

orchestration fee per panel run

multi-model quality and judge output

Provider vault

BYOK control and trace fee

customer credential routing without losing usage governance

Observability broadcast

event fanout meter

send route, credit, fallback, and app events to approved systems

Credit adjustment

coupon, starter credit, or enterprise allocation

separate promotional and paid balances in exports

Endpoint health

serverless or dedicated endpoint test

queue depth, warm cache, quota, and saturation signals

Cost simulation

planning tool included by tier

forecast spend before production route changes

Enterprise policy

governance fee or committed plan

ZDR, region, audit, vendor controls

GPU lane

reserved capacity or usage-based premium

predictable throughput, private supply, and capacity class

Usage examples

Help buyers estimate before they commit.

The page should support concrete planning conversations even before final public rate cards are published.

Use case

Monthly usage

Suggested route

Pricing focus

Support bot

2M input / 600K output

auto-fast

low p95 + cost cap

Code agent

8M input / 2M output

fusion-code

quality + tool use

Research workflow

20M input / 3M output

fusion

coverage + confidence

Private copilot

12M input / 4M output

fusion-private

ZDR + audit trail

Batch automation

100M input / 20M output

value route

margin + throughput

GPU capacity

Compute commitments become a premium pricing path.

Once traffic becomes predictable, Aurona can move customers from public API pass-through into reserved or dedicated GPU-backed routes.

Capacity

Best for

Pricing mode

Serverless endpoint

shared pools for early routes, endpoint tests, and app bursts

usage credits

Burst

short spikes above public provider limits

usage premium

Reserved

predictable production throughput

monthly commitment

Dedicated endpoint

private inference lane for selected models

capacity contract

Sovereign

approved region and vendor policy

custom procurement

Hybrid

public API plus private fallback

blended credits

Buyer paths

One pricing system for demand and supply.

Aurona's pricing can serve developers who spend credits, enterprises who buy commitments, and compute partners who sell inference capacity.

Buyer

Goal

Pricing object

Motion

Developer

test models and routes

starter credits

self-serve

Product team

ship AI feature

usage budgets

pay as you go

AI app company

protect margin

committed credits

growth plan

Enterprise

govern AI usage

procurement + policy

custom plan

Compute partner

sell capacity

GPU lane settlement

capacity agreement

Billing controls

Credits need operational controls around them.

Pricing is strongest when every credit can be traced to a model, route, app, customer, provider, and compute lane.

Credit wallet

Track prepaid credits, committed credits, coupons, invoices, usage by app, and remaining budget from one account.

Bundle planning

Preview credit bundles, top-up behavior, expiration policy, and tier-limit effects as internal planning data before any public offer is approved.

Cost transparency

Separate model cost, Fusion orchestration, policy controls, provider-vault routing, broadcast logs, fallback, and GPU capacity so buyers understand value.

Budget controls

Set credit limits by app, team, customer, environment, route, and provider with alerts and hard stops.

Key-level usage

Let owners inspect each key's spend trend, budget progress, log shortcut, source policy, and rotation status before increasing limits.

Endpoint accounting

Separate serverless endpoint tests, dedicated endpoint usage, warm-cache samples, quota events, and partner capacity lines.

Margin management

Let AI apps apply credit multipliers and route ceilings so customer pricing can remain profitable.

Observability destinations

Manage where usage, route, budget, and app events are sent so finance and operations can reconcile the same credit ledger.

Enterprise commitments

Support annual commitments, volume tiers, dedicated support, SLAs, procurement reporting, and private capacity.

Provider settlement

Use the same ledger to clear model providers, GPU partners, app developers, and Aurona platform fees.

Pricing FAQ

Answer the questions buyers ask before they put real traffic here.

A serious pricing page should explain token billing, route pinning, price changes, fallback attempts, and GPU commitments without waiting for a sales call.

Question

Aurona answer

Operational detail

How are tokens billed?

Model input, model output, Fusion judge, media units, route fees, provider-vault events, broadcast logs, endpoint usage, coupons, and GPU capacity are separated in the credit ledger.

Every line item maps to a model, route, app, customer, environment, and provider.

Can teams pin routes?

Yes. Use exact models for audited behavior or route aliases for managed fallback and migration.

Pinned routes can still keep budget keys, traces, and policy checks.

What happens when pricing changes?

Aurona can surface updated provider prices, enforce route ceilings, and alert before traffic crosses a budget policy.

Enterprise customers can use committed credits and approved-provider lists.

Are failed fallbacks billed?

A production ledger should distinguish successful model tokens, failed provider attempts, retries, and platform fees.

The UI should make retry cost visible before teams scale traffic.

How is cached input shown?

The simulated ledger separates fresh input, provider-reported cache reads, cache writes, and misses instead of blending them into one token total.

Actual credit conversion follows the selected provider and route; this page does not publish a universal cache rate.

How do GPU lanes price?

Serverless endpoint, burst, reserved, dedicated, and hybrid lanes can be priced as usage premiums or monthly capacity commitments.

The same AI Token account settles public model APIs and private inference.

Can teams forecast a route before launch?

Yes. A pricing simulator should estimate model mix, Fusion panel size, fallback behavior, provider key mode, app fees, and GPU service tier.

The forecast is a planning surface, not a public rate-card commitment.

Can credits be planned before purchase?

Yes. The credit bundle planner can compare starter, top-up, committed-credit, and endpoint scenarios without changing public prices.

Finance approval and contract review remain separate from the simulator.

Aurona.ai

Build on the token network for AI applications, models, and compute.

API

OpenAI-compatible

Billing

AI Token credits

Capacity

GPU-backed routes