Developer
Free to start
For prototypes, model testing, and early route integration.
- API keys
- model catalog
- Chat playground
- usage logs
- starter credits
- coupon ledger
Pricing
A developer-friendly path from free testing to AI token credits, production routes, Fusion orchestration, enterprise policy, and GPU-backed capacity.
OpenAI-compatible
Token ledger
GPU route layer
Start
Free
Scale
Credits
Enterprise
Custom
Credit formula
model cost + Fusion orchestration + policy controls + GPU capacity + app margin
transparent model price
credit wallet
budget keys
capacity commitments
Pricing signals
Aurona pricing should separate raw model cost, route intelligence, Fusion quality, enterprise governance, app margin, and GPU capacity.
Pass-through
model cost
Expose public model input and output pricing as the baseline for transparent credit conversion.
Route value
Fusion fee
Price the orchestration layer that selects panels, judges outputs, applies policy, and returns traces.
Margin control
budget keys
Map usage to apps, customers, teams, environments, and route ceilings before traffic scales.
Compute premium
GPU lanes
Plan serverless endpoints, burst pools, reserved lanes, dedicated endpoints, and next-generation GPU capacity as separate commercial paths.
Credits ledger
balance + coupon
Keep paid credits, starter credits, enterprise credits, coupons, and partner settlement lines distinct.
Credit bundles
planner only
Model starter, promotional, top-up, and committed-credit scenarios without publishing a binding rate card.
Key mode
BYOK or managed
Show when customer provider keys, Aurona-managed supply, or hybrid fallback changes the commercial object.
Planning
simulator
Estimate model, Fusion, fallback, app attribution, broadcast logs, and GPU exposure before commitment.
Cache accounting
read + write
Separate provider-reported cache status from fresh input so repeated prompts and tool schemas remain auditable.
Key control
usage tabs
Expose per-key spend charts, budget progress, log shortcuts, and rotation state inside the workspace.
Credit calculator
250,000
Estimated monthly credits
6
model tokens
5
Fusion
2
GPU lane
0
provider vault
0
broadcast logs
0
{
"plan": "Builder",
"messages": 250000,
"model_tokens": "4.50 credits",
"fusion": "1.75 credits",
"gpu_lane": "0.00 credits",
"provider_key_mode": "aurona_managed",
"observability_broadcast": "0.25 credits",
"credit_adjustment": "0.00 credits",
"endpoint_health": "not_selected",
"volume_multiplier": 0.92,
"estimated_total": "6 credits"
}Request reconciliation lab · simulated
Trace provider attempts, usable output, fallback recovery, auxiliary work, and provisional AI Token exposure. This planning view does not create a billing policy or issue credits.
Reconciliation preview
aurona/auto-fast · serverless route
Attempts
1
Usable output
observed
Modeled exposure
0.038 AI Tokens
Provider attempts
approved text pool
usable output
Provisional lines
model input + output
0.032
provisional usage
route evidence
0.006
provisional usage
Simulation only. This preview does not decide charges, issue credits, change billing policy, allocate capacity, or guarantee a refund.
receipt · reconcile:usable-text:fallback-on:0.038
Usage Settlement Lab · simulated
Model an AI Token reservation across text, image, audio, or video, then reconcile only when complete provider-reported usage is present.
Proposed preview
/v1/credits/settlement-preview?workload=text&tier=shared
Capacity
serverless
Receipt
settlement_00vquer1
Modeled reservation
4.02 AI Tokens
Settlement preview
Pending evidence
input
native token meter
1,200 tokens
cached input
native token meter
300 tokens
reasoning
native token meter
180 tokens
output
native token meter
240 tokens
01
request estimated
02
AI Token reservation modeled
03
provider usage pending
04
wallet unchanged
Simulated translation weights and usage only. This lab does not debit AI Tokens, change a wallet, publish a rate card, or reserve GPU capacity.
Credit Headroom Lab · simulated
Join AI Token balance, a workspace reserve, in-flight demand, and modeled GPU-backed route headroom in one admission receipt.
Planning receipt
headroom_b4565b26
POST /v1/credits/headroom-preview
Spendable
100 AI Tokens
Request estimate
12 AI Tokens
Lane headroom
4 after preview
AI Token reserve and modeled lane headroom both remain reviewable.
Modeled ceiling: 18 · current in flight: 8 · requested: 6
01
balance evidence reviewed
02
workspace reserve preserved
03
request exposure modeled
04
lane admission previewed
Demonstration values only. This preview does not debit AI Tokens. It does not set a production rate limit, promise throughput, change billing, or allocate GPU capacity.
Plans
Aurona can support early developers, production teams, AI app companies, enterprises, and compute partners without changing the underlying API.
Developer
For prototypes, model testing, and early route integration.
Builder
For teams shipping production traffic with transparent model costs.
Growth
For apps that need predictable volume, margin control, and support.
Enterprise
For governed usage, procurement, private policy, and reserved capacity.
Workspace billing
A user starts with a personal workspace, then upgrades into a shared company or enterprise workspace when multiple people need the same keys, budgets, logs, data policies, and invoices.
Workspace
Billing basis
What it manages
Commercial motion
Personal workspace
Free sandbox + starter credits
one owner, project API keys, playground, usage logs
upgrade to Company when a team needs shared billing
Company workspace
Pay-as-you-go AI Token credits
members, roles, projects, budgets, invoices, customer attribution
set monthly caps and route-level hard stops
Growth workspace
Committed credits
volume planning, app margins, priority routes, usage reporting
reserve credits for predictable production traffic
Enterprise workspace
Custom contract and invoice
SSO, audit logs, ZDR, data policy, procurement, SLAs
attach private routes or reserved GPU capacity
Compute partner workspace
Capacity settlement
GPU inventory, SLA, DD evidence, facility data, network profile
settle approved supply through capacity agreements
Credit meters
Different AI workloads can settle through one wallet while preserving model-specific, route-specific, and compute-specific economics.
Workload
Meters
Pricing drivers
Model tokens
input + output tokens
public model price, provider, context, volume
Cached input
cache reads + cache writes
provider-reported cache status, route affinity, reusable context, and miss reason
Starter and promo credits
coupon, trial, or enterprise allocation
expiration, workspace, and invoice treatment
Credit bundle scenarios
starter, top-up, committed, or partner credit plans
non-binding planning view before billing approval
Fusion routes
panel + judge + final tokens
panel size, judge model, trace level
Multimodal apps
media units + tokens
image/audio/video provider and output size
Tool and agent runs
tool calls + tokens + retries
workflow depth and fallback policy
Provider vault
customer-owned provider key usage
key mode, fallback, trace, and route margin
API key controls
per-key usage, budget, and log filters
workspace, route, environment, allowlist, and rotation state
Tier limit previews
request volume, token volume, modality, endpoint mode
account governance and launch review
Broadcast logs
trace fanout events
logs, webhooks, billing exports, analytics sinks
Serverless endpoints
endpoint seconds, warm-cache tests, queue samples
GPU class, region, quota, and runtime mode
Dedicated inference
reserved capacity + usage
GPU class, region, throughput, commitment, endpoint mode
Commercial layers
A great pricing page should show exactly what Aurona is monetizing: model access, routing intelligence, Fusion quality, policy governance, and GPU-backed capacity.
Layer
Pricing basis
Value delivered
Public model access
pass-through model cost + route margin
transparent source pricing
Aurona Auto
routing and fallback fee
cost, latency, policy, provider health
Aurona Fusion
orchestration fee per panel run
multi-model quality and judge output
Provider vault
BYOK control and trace fee
customer credential routing without losing usage governance
Observability broadcast
event fanout meter
send route, credit, fallback, and app events to approved systems
Credit adjustment
coupon, starter credit, or enterprise allocation
separate promotional and paid balances in exports
Endpoint health
serverless or dedicated endpoint test
queue depth, warm cache, quota, and saturation signals
Cost simulation
planning tool included by tier
forecast spend before production route changes
Enterprise policy
governance fee or committed plan
ZDR, region, audit, vendor controls
GPU lane
reserved capacity or usage-based premium
predictable throughput, private supply, and capacity class
Usage examples
The page should support concrete planning conversations even before final public rate cards are published.
Use case
Monthly usage
Suggested route
Pricing focus
Support bot
2M input / 600K output
auto-fast
low p95 + cost cap
Code agent
8M input / 2M output
fusion-code
quality + tool use
Research workflow
20M input / 3M output
fusion
coverage + confidence
Private copilot
12M input / 4M output
fusion-private
ZDR + audit trail
Batch automation
100M input / 20M output
value route
margin + throughput
GPU capacity
Once traffic becomes predictable, Aurona can move customers from public API pass-through into reserved or dedicated GPU-backed routes.
Capacity
Best for
Pricing mode
Serverless endpoint
shared pools for early routes, endpoint tests, and app bursts
usage credits
Burst
short spikes above public provider limits
usage premium
Reserved
predictable production throughput
monthly commitment
Dedicated endpoint
private inference lane for selected models
capacity contract
Sovereign
approved region and vendor policy
custom procurement
Hybrid
public API plus private fallback
blended credits
Buyer paths
Aurona's pricing can serve developers who spend credits, enterprises who buy commitments, and compute partners who sell inference capacity.
Buyer
Goal
Pricing object
Motion
Developer
test models and routes
starter credits
self-serve
Product team
ship AI feature
usage budgets
pay as you go
AI app company
protect margin
committed credits
growth plan
Enterprise
govern AI usage
procurement + policy
custom plan
Compute partner
sell capacity
GPU lane settlement
capacity agreement
Billing controls
Pricing is strongest when every credit can be traced to a model, route, app, customer, provider, and compute lane.
Track prepaid credits, committed credits, coupons, invoices, usage by app, and remaining budget from one account.
Preview credit bundles, top-up behavior, expiration policy, and tier-limit effects as internal planning data before any public offer is approved.
Separate model cost, Fusion orchestration, policy controls, provider-vault routing, broadcast logs, fallback, and GPU capacity so buyers understand value.
Set credit limits by app, team, customer, environment, route, and provider with alerts and hard stops.
Let owners inspect each key's spend trend, budget progress, log shortcut, source policy, and rotation status before increasing limits.
Separate serverless endpoint tests, dedicated endpoint usage, warm-cache samples, quota events, and partner capacity lines.
Let AI apps apply credit multipliers and route ceilings so customer pricing can remain profitable.
Manage where usage, route, budget, and app events are sent so finance and operations can reconcile the same credit ledger.
Support annual commitments, volume tiers, dedicated support, SLAs, procurement reporting, and private capacity.
Use the same ledger to clear model providers, GPU partners, app developers, and Aurona platform fees.
Pricing FAQ
A serious pricing page should explain token billing, route pinning, price changes, fallback attempts, and GPU commitments without waiting for a sales call.
Question
Aurona answer
Operational detail
How are tokens billed?
Model input, model output, Fusion judge, media units, route fees, provider-vault events, broadcast logs, endpoint usage, coupons, and GPU capacity are separated in the credit ledger.
Every line item maps to a model, route, app, customer, environment, and provider.
Can teams pin routes?
Yes. Use exact models for audited behavior or route aliases for managed fallback and migration.
Pinned routes can still keep budget keys, traces, and policy checks.
What happens when pricing changes?
Aurona can surface updated provider prices, enforce route ceilings, and alert before traffic crosses a budget policy.
Enterprise customers can use committed credits and approved-provider lists.
Are failed fallbacks billed?
A production ledger should distinguish successful model tokens, failed provider attempts, retries, and platform fees.
The UI should make retry cost visible before teams scale traffic.
How is cached input shown?
The simulated ledger separates fresh input, provider-reported cache reads, cache writes, and misses instead of blending them into one token total.
Actual credit conversion follows the selected provider and route; this page does not publish a universal cache rate.
How do GPU lanes price?
Serverless endpoint, burst, reserved, dedicated, and hybrid lanes can be priced as usage premiums or monthly capacity commitments.
The same AI Token account settles public model APIs and private inference.
Can teams forecast a route before launch?
Yes. A pricing simulator should estimate model mix, Fusion panel size, fallback behavior, provider key mode, app fees, and GPU service tier.
The forecast is a planning surface, not a public rate-card commitment.
Can credits be planned before purchase?
Yes. The credit bundle planner can compare starter, top-up, committed-credit, and endpoint scenarios without changing public prices.
Finance approval and contract review remain separate from the simulator.
Aurona.ai
API
OpenAI-compatible
Billing
AI Token credits
Capacity
GPU-backed routes