Quickstart
Swap your base URL, add an Aurona API key, choose aurona/auto, and send your first request.
Developer documentation
Start with one OpenAI-compatible request, then add intelligent routing, AI Token controls, policy, and GPU capacity.
First routed request
TypeScript · OpenAI-compatible
const response = await fetch("https://api.aurona.ai/v1/chat/completions", {
method: "POST",
headers: { Authorization: `Bearer ${AURONA_API_KEY}` },
body: JSON.stringify({
model: "aurona/auto",
route: { objective: "balance" },
messages: [{ role: "user", content: "Build an agent." }]
})
});Developer signals
The strongest API docs show both the familiar starting point and the differentiated controls: routes, credits, policy, GPU lanes, traces, and events.
Compatibility
OpenAI SDK
Start by changing base URL and API key, then add route, budget, policy, and trace fields.
Route config
JSON object
Document provider sort, fallback, data policy, service tier, budget key, compute lane, and return trace.
Ledger events
credits
Expose model, Fusion, media, GPU, app, and platform fee events for billing transparency.
Operations
broadcast
Send route traces into webhooks, logs, observability pipelines, billing exports, and app dashboards.
Account console
keys + credits
Document workspace menus for scoped keys, credits, coupons, usage reports, endpoint state, and labs.
Agent connector
docs + live data
Let coding assistants inspect model catalog, route prices, credit state, rankings, and safe test calls through a scoped connector.
Agent policies
route + tools
Bound coordinator and worker routes, allowed tools, child tasks, credit exposure, and compute preference before a run.
Budget windows
daily to lifetime
Show how workspace budgets can enforce daily, weekly, monthly, and lifetime credit ceilings before requests run.
Integration paths
Start with a direct request, preserve an existing OpenAI SDK integration, or give a scoped agent live catalog and docs context. Every path converges on the same Aurona routes, AI Token ledger, workspace policy, and GPU-backed capacity model.
Path
Starting point
Aurona control
Best use
Direct REST
Any server runtime; no additional client dependency
Full request control for exact models, Aurona routes, streaming, tools, and response metadata
Teams building a custom API layer or testing the first request
OpenAI SDK
Existing OpenAI-compatible application
Swap the base URL and key, then add Aurona route, budget, policy, and trace fields when needed
Teams migrating an established chat, agent, or tool-calling integration
Agent connector
Scoped coding tool or internal agent session
Read catalog, rankings, docs, and credit state; keep metered tests explicit and bounded
Builders who need live platform context while they implement or review a route
Project bootstrap lab · simulated
Plan the integration path, scoped access, masked environment manifest, AI Token ceiling, live build context, verification steps, and GPU-capacity handoff before any production action.
Bootstrap preview
OpenAI-compatible API · Personal sandbox
Key ownership
Personal project key plan
Modeled use
6 / 12 AI Tokens
Capacity
Serverless first
Masked environment manifest
Live build context
Verification and handoff
Start with a managed route and watch workload evidence
Simulation only: this packet does not create or rotate credentials, does not write environment files, does not enable billing, does not allocate GPU capacity, and does not change production traffic.
receipt · bootstrap-preview:api-app:personal-sandbox:serverless-first:12
Route onboarding
This simulated planning path keeps early integration simple while making the handoff to route review, workspace controls, and GPU-backed capacity legible.
Step
What the builder does
What Aurona keeps visible
Request builder
Send one OpenAI-compatible test with an exact model or aurona/auto.
Save the prompt, parameters, and credit estimate as a simulated planning path.
Route evidence
Compare catalog fit, ranking signals, provider policy, fallback, and budget scope.
Keep the selected route explainable before it receives app or production traffic.
Capacity promotion
Move a validated workload through serverless, reserved, or dedicated endpoint choices.
Review modeled endpoint health, ownership, and budget context; no capacity commitment is created here.
Model discovery
Filter model supply by output type and supported request controls, then sort by economics, context, performance, demand, or catalog age. Workspace aliases can follow an approved upgrade policy while production routes remain version-pinned.
Catalog query
const models = await fetch(
"https://api.aurona.ai/v1/models?output_modalities=text,image&supported_parameters=tools&sort=throughput-high-to-low",
{ headers: { Authorization: `Bearer ${AURONA_API_KEY}` } }
).then((response) => response.json());
// Use a workspace alias for controlled upgrades; pin exact versions in production.
const model = "aurona/latest-balanced";Catalog query preview
The proposed Aurona preview intersects hard capability, capacity, provider-declared data use, and access requirements before sorting visible evidence or resolving a workspace alias. It returns excluded reasons and a stable planning-only receipt. It does not change aliases or policies, send traffic, debit AI Tokens, publish models, or allocate GPU capacity.
Evidence
Returned fields
Planning use
Hard filters
output modality, required parameter, approved capacity
Exclude incompatible supply before any sort runs
Provider-declared conditions
data-use class and access class
Fail closed or route supply into workspace review before alias resolution
Excluded evidence
canonical ID and review reason
Show why otherwise-capable supply was removed before scoring
Sort evidence
AI Token estimate, throughput, latency, or demand
Place unmeasured candidates last instead of treating missing data as zero
Canonical identity
workspace alias → immutable route version
Keep discovery convenient while production evidence remains attributable
Lifecycle
preview, current, maintenance, retiring
Show maturity and replacement context alongside capability data
Planning receipt
normalized query + eligible canonical IDs
Make the same read-only query reproducible across workspace reviews
Capacity path
serverless or dedicated review
Carry model demand toward GPU planning without allocating capacity
Catalog query plan
const preview = await fetch(
"https://api.aurona.ai/v1/models/query-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
output_modality: "text",
required_parameter: "tools",
sort: "throughput_high_to_low",
alias_policy: "workspace_alias",
capacity: "serverless",
data_use_policy: "no_contribution_only",
access_policy: "standard_only",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed read-only response: eligible count, canonical route IDs,
// provider-declared data use, access review, exclusions,
// AI Token estimate, capacity path, and planning-only receipt.Variant contract preview
The proposed Aurona preview resolves an immutable simulated variant only after its delivery mode, reasoning class, native meter, task contract, modeled AI Token amount, and capacity path are declared. A missing or incompatible field holds the alias instead of silently substituting a different execution contract.
Evidence
Returned fields
Planning use
Family identity
stable family plus distinct variant ID
Keep discovery convenient without flattening operationally different routes
Delivery mode
interactive request or deferred job
Fail closed before a synchronous client receives an asynchronous contract
Decision shape
yes/no, allowed choice, or bounded score with abstention policy
Keep typed decision workloads distinct from generated prose and verify schema fit before routing
Reasoning class
standard or deliberate review
Keep higher-reasoning supply and its meter explicit
Native meter
input, output, reasoning, character, image, second, or request
Preserve the source unit before modeled AI Token translation
Task contract
synchronous response or task ID plus polling
Make recovery and status handling visible before integration
Capacity path
serverless, shared batch, or reserved review
Join execution mode to GPU planning without allocating capacity
Planning receipt
normalized request and selected immutable variant
Make alias resolution reviewable and reproducible
Variant contract plan
const preview = await fetch(
"https://api.aurona.ai/v1/models/variant-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
family: "aurona/atlas",
delivery_mode: "deferred",
reasoning_class: "standard",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning-only response: immutable variant ID, native meter,
// synchronous or asynchronous task contract, modeled AI Tokens,
// capacity path, exclusions, and deterministic receipt.Protocol route preview
The proposed Aurona preview evaluates Chat Completions, Responses, Anthropic Messages, and asynchronous media jobs as explicit route inputs. Native support ranks first; reviewed adapters remain visible; approved fallback attempts share one request contract. Any feature, adapter, capacity, or fallback loss produces a hold.
Evidence
Returned fields
Planning use
Incoming contract
Chat Completions, Responses, Anthropic Messages, or asynchronous video job
Resolve the route from the protocol the application already uses
Hard feature
tools, structured output, reasoning, streaming, media references, or async result
Exclude any path that cannot preserve the required behavior
Translation state
native or reviewed adapter
Prefer native support and make approved conversion visible
Fail-closed hold
protocol, feature, adapter, or capacity mismatch
Never silently relax the application contract
Commercial evidence
modeled AI Token estimate
Compare compatible paths without publishing a rate card
Capacity path
serverless or dedicated review
Join protocol compatibility to GPU planning without allocating capacity
Fallback contract
disabled or approved chain; shared request fields across attempts
Reject per-attempt overrides that could silently change application behavior
Ordered attempts
primary, fallback-1, fallback-2, trigger reasons
Hold when no second route preserves every protocol, feature, policy, and capacity gate
Protocol compatibility plan
const preview = await fetch(
"https://api.aurona.ai/v1/routes/protocol-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
protocol: "anthropic_messages",
required_feature: "tools",
adapter_policy: "reviewed_adapter",
capacity: "all",
fallback_policy: "approved_chain",
attempt_policy: "shared_contract",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning response: native and reviewed-adapter candidates,
// ordered fallback roles, shared-contract evidence, exclusions,
// modeled AI Token evidence, capacity, and receipt.Ranking feed preview
The proposed Aurona feed returns simulated model, route, and app rows with a declared measure, freshness, version, schema, canonical endpoint, portable citation, visibility scope, AI Token context, and GPU capacity path. Public previews fail closed on workspace-only evidence, while provenance metadata makes internal exports attributable without asserting reuse rights.
Evidence
Returned fields
Planning use
Dataset identity
as-of timestamp, version, schema, format, canonical endpoint
Keep every exported ranking snapshot attributable and reproducible
Portable citation
publisher, title, version, measure, window, scope, simulated origin
Carry human-readable provenance into internal reports without implying reuse rights
Measure semantics
modeled AI Token volume or request-share index
Separate adoption evidence from quality, latency, and benchmark claims
Visibility scope
workspace view or public opt-in only
Exclude workspace-only rows before ranking a public preview
Surface filter
models, routes, apps, or all eligible
Return only the operating surface requested by the consumer
Ranking rows
absolute rank, stable id, modeled requests
Preserve ordering after every hard filter
Capacity evidence
serverless, reserved review, dedicated review
Join demand signals to GPU planning without allocating capacity
Ranking feed query
const feed = await fetch(
"https://api.aurona.ai/v1/datasets/rankings-preview?surface=routes&scope=workspace-only&metric=ai-token-volume&window=30d",
{ headers: { Authorization: `Bearer ${AURONA_API_KEY}` } }
).then((response) => response.json());
// Proposed planning response: ranked rows plus as_of, version, resolved window,
// dataset manifest, portable citation, measure semantics, visibility scope,
// capacity evidence, and a stable receipt.
// It does not publish traffic, change visibility, debit credits, or reserve GPUs.Native meter ranking preview
The proposed Aurona preview ranks simulated routes only when measure, native unit, completed window, and visibility scope match. Tokens, images, tool calls, media time, and GPU hours remain separate. A within-track index shows relative position without inventing a cross-unit score.
Evidence
Returned fields
Planning use
Track contract
measure, native unit, completed window
Keep unlike usage evidence out of the same ordering
Native units
tokens, images, tool calls, audio seconds, video seconds, GPU hours
Preserve how each product surface actually measures work
Within-track index
leader equals 100; every other row is relative to that leader
Compare position without inventing cross-unit conversion
Visibility gate
public opt-in or workspace review
Suppress private evidence before a public ranking is calculated
Freshness
latest completed evidence bucket in UTC
Distinguish data freshness from page-render time
Capacity evidence
serverless, reserved review, dedicated review
Connect demand to GPU planning without allocating capacity
Planning receipt
normalized track query and eligible stable IDs
Make the same read-only ordering reproducible
Native meter ranking plan
const preview = await fetch(
"https://api.aurona.ai/v1/rankings/native-meter-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
measure: "tool-activity",
native_unit: "tool-calls",
window: "7d",
visibility: "public",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning-only response: ranked rows in one native unit,
// a within-track index, completed-window freshness, visibility and
// capacity evidence, exclusions, and a deterministic receipt.Multimodal evaluation preview
The proposed Aurona preview compares simulated image routes against the same prompt family and evaluation floor, then ranks only eligible supply by quality, generation speed, or modeled AI Token efficiency. Capacity mismatches fail closed before a route is selected.
Evidence
Returned fields
Planning use
Shared prompt family
exact text, editing, consistency, counting, spatial
Compare candidates against the same observable capability
Evidence floor
case count and evaluation version
Keep small or stale samples out of route selection
Objective
quality, generation speed, or AI Token efficiency
Declare the ranking goal instead of hiding a blended score
Native evidence
pass rate, seconds, modeled AI Tokens
Preserve each measure behind the selected route
Capacity scope
serverless, reserved review, dedicated review
Fail closed when the requested GPU path has no eligible route
Planning receipt
normalized gates and eligible route IDs
Make a read-only comparison repeatable before production review
Multimodal route plan
const preview = await fetch(
"https://api.aurona.ai/v1/evaluations/multimodal-route-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
modality: "image",
prompt_family: "exact-text",
objective: "quality",
minimum_cases: 12,
capacity: "reserved-review",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning response: eligible and excluded routes, native evidence,
// capacity scope, selected route preview, evaluation version, and receipt.Model alias lifecycle
Aurona aliases separate discovery from release control. Moving and preview aliases reduce catalog churn, while pinned versions and review evidence keep production routes attributable across policy, credits, evaluations, and capacity.
State
Example
Control
Best use
Moving alias
aurona/latest-balanced
Follows a workspace-approved family policy
Exploration and controlled development environments
Preview alias
aurona/preview-balanced
Exposes the next candidate to saved evaluations, policy checks, and cost review
Pre-production comparison without changing the stable route
Pinned version
aurona/balanced-2026-08
Keeps model, provider requirements, parameters, and capacity class attributable
Production routes that need deliberate releases and repeatable evidence
Retiring alias
aurona/balanced-2026-05
Shows replacement route, owner, evaluation state, and retirement window
Planned migration with logs and fallback evidence
Review before promotion
preview → stable
Checks workload quality, tools, AI Token budget, workspace policy, and GPU lane fit
A simulated review that does not move production traffic
API reference
Aurona keeps the first request familiar while adding the controls AI companies need as they scale.
Advanced routed request
const response = await fetch(
"https://api.aurona.ai/v1/chat/completions",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "aurona/auto",
route: {
objective: "balance",
optimize: ["price", "latency", "policy"],
cost_quality_tradeoff: 6,
fallback: "auto",
session_id: "agent-thread-42",
data_policy: "zero_retention",
service_tier: "serverless",
endpoint_type: "managed_pool",
budget_key: "prod-agent",
budget_interval: "monthly",
policy_rules: ["mask-sensitive-metadata", "approved-provider-order"],
ranking_scope: "route",
app_id: "internal-agent-workbench"
},
usage_filter: {
timezone: "workspace",
granularity: "hourly"
},
playground_history: {
save: true,
redaction: "metadata_only"
},
messages: [{ role: "user", content: "Build an agent." }]
})
}
);Endpoint
Use
Layer
Area
/v1/chat/completions
Chat, agents, tools, streaming
OpenAI-compatible
core
/v1/models
Model and route catalog
provider-aware
catalog
/v1/models?output_modalities=
Filter text, image, audio, and embedding outputs
catalog
discovery
/v1/models/query-preview
Preview capabilities, provider-declared supply conditions, canonical identity, credits, and capacity
catalog
planning
/v1/models/variant-preview
Preview delivery mode, reasoning class, native meter, task contract, credits, and capacity
catalog
planning
/v1/routes/protocol-preview
Preview native or reviewed-adapter protocol support plus a contract-safe fallback chain
routing
planning
/v1/models/capabilities
Capability, region, training-review, and parameter filters
catalog
discovery
/v1/models/lifecycle
Release state, version policy, retirement date, and replacement route
catalog
operations
/v1/fusion/runs
Panel, judge, synthesis
multi-model
quality
/v1/routes
Create and manage route aliases
routing
control
/v1/routes/decide
Simulate task classification, approved-pool scoring, fallback, and decision receipt
routing
planning
/v1/routes/session-continuity-preview
Preview session route reuse, switching, policy holds, budgets, and GPU lane evidence
routing
planning
/v1/routes/tool-evidence-preview
Rank tool-capable endpoints with schema, evaluation, throughput, credit, and capacity evidence
routing
planning
/v1/routes/eligibility-preview
Preview layered workspace, member, key, budget, regional execution, retention, and capacity eligibility
routing
governance
/v1/route-objectives
Reusable cost, quality, latency, balance, and policy objective presets
routing
control
/v1/routes/:id/metadata
Route evidence, provider choice, policy action, credit trace
routing
trace
/v1/management/keys
Organization and automation keys
admin
control
/v1/api-keys/:id/usage
Per-key usage, budgets, and spend charts
account
control
/v1/api-keys/:id/logs
Pre-filtered logs for one API key
account
trace
/v1/mcp
Live model catalog and route testing for agents
tool server
developer
/v1/agent-connectors
Scoped connector sessions for coding tools and internal agents
developer
setup
/v1/projects/bootstrap-preview
Simulate integration, scoped access, AI Token budget, and capacity handoff
developer
planning
/v1/projects/context-manifest
Return a masked environment and live-context manifest without issuing credentials
developer
planning
/v1/agent-policies
Coordinator, worker, tool, budget, and capacity controls
agent runtime
control
/v1/workload-labels
Workspace label dimensions, sampling, and usage rollups
observability
governance
/v1/cli/sessions
CLI login, profile, and route bootstrap
developer
setup
/v1/provider-keys
Customer-owned provider key vault
routing
privacy
/v1/provider-keys/route-preview
Preview masked credential scope, health, priority, fallback, and capacity
routing
planning
/v1/model-quickstarts
Copy-ready model examples by modality
docs
onboarding
/v1/model-shortlists
Compare evidence, policy, lifecycle, cost, latency, and GPU readiness
catalog
planning
/v1/evaluations
Create application-specific prompt and tool-assertion suites
evaluation
planning
/v1/evaluations/:id/runs
Run pinned route candidates against the same workload cases
evaluation
evidence
/v1/evaluation-gates
Preview quality, tool, policy, credit, and capacity promotion gates
routing
control
/v1/rankings
Model, route, app, and GPU-ready ranking feeds
market data
rankings
/v1/rankings/native-meter-preview
Rank one declared measure, native unit, window, and visibility scope at a time
market data
planning
/v1/datasets/rankings-preview
Preview versioned ranking rows, measure semantics, visibility scope, citation metadata, and capacity evidence
market data
planning
/v1/evaluations/multimodal-route-preview
Compare shared prompt-family evidence before previewing an image route
evaluation
planning
/v1/policy/rules
Priority rules for prompt, response, provider, route, and budget actions
enterprise
policy
/v1/policy/actions
Block, mask, warn, reroute, or hold requests for review
enterprise
policy
/v1/credits
Wallet, ledgers, balances
token settlement
billing
/v1/credits/settlement-preview
Translate native usage meters into an AI Token reservation and reconciliation preview
token settlement
planning
/v1/credits/headroom-preview
Preview request admission from AI Token reserve, in-flight demand, and route capacity evidence
token control
planning
/v1/budgets
Limits, alerts, hard stops
spend control
finance
/v1/workspaces/:id/budgets
Workspace budget intervals and enforcement state
spend control
finance
/v1/batches
Plan and submit asynchronous text or embedding work
inference
batch
/v1/batches/:id
Inspect validation, queue, execution, reconciliation, and result state
operations
batch
/v1/cost-simulations
Estimate route and model spend
planning
finance
/v1/usage/reconciliation-preview
Simulate delivery outcome, attempts, fallback, and provisional AI Token lines
planning
finance
/v1/usage/requests/:id
Inspect request outcome, route evidence, attempts, and reconciled usage state
operations
trace
/v1/activity/investigation-preview
Preview aggregate queries, prior-window comparison, saved-view scope, and exact request-log filters
observability
planning
/v1/apps
App attribution and metering
ecosystem
apps
/v1/apps/installs
Install, approve, and meter marketplace apps
ecosystem
apps
/v1/apps/submissions
Builder listings, runtime choices, review state
ecosystem
apps
/v1/apps/runtimes
Hosted, self-hosted, serverless, and dedicated app runtime choices
ecosystem
apps
/v1/apps/reviews
Runtime, policy, meter, and playground evidence packets
ecosystem
review
/v1/apps/readiness-preview
Simulate route, meter, policy, evidence, and runtime readiness
ecosystem
planning
/v1/apps/runtime-preview
Preview workspace file scope, sandbox continuity, network policy, routes, credits, and capacity
ecosystem
planning
/v1/apps/launch-packets
Save a reviewed app launch packet without publishing it
ecosystem
review
/v1/workspaces
Members, roles, projects, keys
account
admin
/v1/workspaces/:id/change-preview
Preview affected keys, apps, routes, budgets, and capacity before assignment
account
planning
/v1/workspaces/:id/menu
Workspace switcher, credits, labs, logs, and route shortcuts
account
console
/v1/console/surfaces
Dashboard, model hub, playground, endpoints, storage, and runtime shortcuts
account
console
/v1/billing/exports
Credit, invoice, coupon, app, and partner settlement exports
billing
finance
/v1/policies
Retention, region, vendor rules
enterprise
policy
/v1/guardrails
Prompt, output, and sensitive data checks
enterprise
safety
/v1/gpu-lanes
Reserved and dedicated capacity
compute
capacity
/v1/capacity/promotion-preview
Compare completed demand, utilization, traffic shape, runtime fit, region evidence, and normalized exposure
compute
planning
/v1/inference/endpoints
Serverless and dedicated endpoint inventory
compute
capacity
/v1/inference/endpoints/:id/health
Queue depth, warm cache, quota, and saturation
compute
routing
/v1/inference/endpoints/:id/promotion
Route evidence, health sample, policy, and budget handoff preview
compute
planning
/v1/inference/usage
Serverless and dedicated usage by model, key, region, and time window
compute
billing
/v1/playground/history
Prompt tests, multimodal runs, route decisions, and cost traces
developer
console
/v1/observability/destinations
Manage log, trace, and warehouse sinks
operations
control
/v1/observability/events
Route decisions and usage logs
operations
trace
/v1/observability/broadcasts
Fan out route traces
operations
broadcast
API tester
This tester changes the route and generates the exact response shape developers expect.
Request
TypeScript{
"model": "aurona/auto",
"messages": [{ "role": "user", "content": "Compare provider routes." }],
"provider": {
"only": ["approved-frontier", "approved-fallback"],
"require_parameters": true
},
"route": {
"data_policy": "zero_retention",
"service_tier": "priority",
"budget_key": "docs-demo",
"return_trace": true
},
"observability": {
"broadcast": ["usage_logs", "webhook", "billing_export"]
},
"guardrails": {
"prompt_injection": "block",
"sensitive_data": "metadata_only"
}
}Response
idle{
"id": "chatcmpl_demo_aurona_auto",
"model": "aurona/auto",
"sdk": "TypeScript",
"policy": "zero_retention",
"credits": "0.184",
"latency": "640ms",
"service_tier": "priority",
"broadcast": "queued",
"trace": "returned"
}Authentication
A serious AI API needs more than one secret. Aurona should make environment scoping, budget attribution, and webhook verification explicit.
Key type
Where used
Purpose
Bearer token
Authorization header
server-side API calls
Management key
admin API and automation
provision projects, budgets, routes, and logs
Project key
app or environment scope
separate dev, staging, prod
Route key
alias and budget scope
share one route without exposing the full workspace
Endpoint key
serverless or dedicated endpoint scope
test and meter GPU-backed inference lanes
Policy key
policy rules and review exports
manage rule priority, actions, and evidence packets
Provider key
encrypted vault entry
route approved calls through customer-owned credentials
Budget key
usage and spend scope
customer, team, app, route
Webhook secret
signature verification
billing and route events
Provider keys
Some customers want to bring an existing provider account. Aurona can still own the model catalog, policy checks, budget keys, traces, app attribution, and fallback rules.
Mode
Credential source
Best use
Aurona-managed
Aurona provider account
default model access, credits, route margin
Customer vault
workspace encrypted key
BYOK routing with Aurona policy, logs, and budget controls
Hybrid
Aurona plus customer key
use customer-preferred provider first, keep approved fallback
Private lane
dedicated endpoint or GPU lane
enterprise or compute-partner capacity with token settlement
Provider-key route
{
"model": "aurona/auto",
"provider": {
"key_mode": "customer_vault",
"vault_key": "pv_customer_openai",
"order": ["preferred-frontier", "approved-fallback"],
"only": ["preferred-frontier", "approved-fallback"],
"require_parameters": true,
"allow_fallbacks": true
},
"route": {
"service_tier": "priority",
"budget_key": "workspace-prod",
"cost_quality_tradeoff": 5,
"return_trace": true
},
"observability": {
"destination": "workspace-usage-warehouse",
"logs_shortcut": true
}
}Provider vault route preview
The proposed Aurona contract resolves masked workspace scope, credential readiness, priority, approved fallback, AI Token context, and GPU capacity into a deterministic planning receipt. It never accepts or returns raw provider credentials.
Layer
Evidence
Planning use
Scope
workspace, project key, route, model
Exclude credentials before route ordering
Credential evidence
masked identity, enabled state, expiry, rate-limit state
Show whether customer supply is usable without returning a secret
Priority
customer first, managed first, customer only
Make the attempt order explicit
Fallback
approved managed supply or fail-closed hold
Keep policy behavior visible when customer supply is unavailable
Capacity
serverless pool or reserved review
Connect the route plan to GPU-backed inference capacity
Receipt
scope, attempts, reason, AI Token context
Return deterministic planning evidence without changing traffic
Provider vault route preview
const preview = await fetch(
"https://api.aurona.ai/v1/provider-keys/route-preview",
{
method: "POST",
headers: {
Authorization: "Bearer <management-key>",
"Content-Type": "application/json"
},
body: JSON.stringify({
workspace_id: "workspace-aurona-launch",
project_key_id: "ak_prod_masked",
route: "aurona/auto",
key_mode: "customer_first",
credential_ref: "pv_masked_7Q2",
fallback_allowed: true,
capacity: "serverless"
})
}
);
// Proposed planning-only contract. It returns masked scope, health, priority,
// fallback, capacity, stages, and a stable receipt. It does not create, inspect, rotate, or delete a credential;
// route live traffic; debit AI Tokens;
// or allocate GPU capacity.Fusion route config
Developers can keep one API call while product, finance, and security teams tune route behavior behind the scenes.
Fusion object
{
"model": "aurona/fusion",
"fusion": {
"panel": [
"openai/gpt-5.5-pro",
"anthropic/claude-sonnet-4.6",
"google/gemini-3.5-flash"
],
"judge": "aurona/judge-premium",
"return_analysis": true
},
"provider": {
"sort": "throughput",
"allow_fallbacks": true,
"data_collection": "deny",
"cost_quality_tradeoff": 7,
"session_id": "fusion-review-2026-07"
}
}Route options
The docs should make token credits, policy routing, provider selection, and GPU capacity feel concrete.
Field
Values
Purpose
model
aurona/auto, aurona/fusion, exact model id
Selects the route or model
route.objective
cost, quality, latency, balance, policy
Chooses a reusable routing objective before provider scoring
route.allowed_pool
workspace-approved, workspace-private-gpu
Limits model selection before scoring and fallback
route.return_decision_receipt
true, false
Returns workload class, objective, eligible pool, selected route, fallback, capacity, and simulated estimate
provider.sort
price, throughput, latency
Chooses provider objective
catalog.sort
price, context, throughput, latency, popularity, newest
Orders discovery results before route evaluation
catalog.output_modalities
text, image, audio, embeddings
Keeps multimodal discovery on one catalog surface
catalog.supported_parameters
tools, reasoning, structured_outputs
Returns only models that can honor required request controls
catalog.data_use_policy
no_contribution_only or provider_declared
Keeps provider-declared data use inside explicit workspace review
catalog.access_policy
standard_only or review_required
Fails closed on gated supply before an alias resolves
protocol.input
chat_completions, responses, anthropic_messages, video_job
Pins the incoming request contract before route selection
protocol.adapter_policy
preserve_only, reviewed_adapter
Holds the request when translation would lose a required feature
evaluation.suite_id
workspace evaluation id
Pins every route candidate to the same application cases and assertions
evaluation.release_gate
quality, tools, policy, credits, latency
Returns review evidence before a route alias or GPU lane is promoted
provider.order
approved provider IDs
Sets preferred provider order
provider.only
approved provider IDs
Limits a request to a specific vendor set
provider.key_mode
aurona, customer_vault, hybrid
Chooses whose provider credential funds the request
provider.require_parameters
true, false
Skips providers that cannot honor requested tools or schemas
allow_fallbacks
true, false
Controls substitution on failure
cost_quality_tradeoff
0-10
Tunes the route between cheapest healthy model and highest-quality answer
session_id
string
Keeps multi-turn agent work sticky to a compatible provider when useful
session.continuity
reuse, switch, hold
Returns whether the next turn can keep its route after task, health, policy, AI Token, and capacity checks
session.capacity_preference
serverless, reserved
Lets continuity planning include a GPU lane review without allocating capacity
data_policy
zero_retention, regional, private
Controls prompt handling
service_tier
shared, priority, serverless, reserved, dedicated
Chooses throughput and capacity class
compute_lane
serverless, hosted, gpu-fast, dedicated
Controls capacity path
endpoint.type
serverless, dedicated, private
Chooses managed pool, tenant lane, or private endpoint
endpoint.health_policy
queue_depth, warm_cache, quota, saturation
Lets the router avoid unhealthy capacity
endpoint.runtime
vLLM, TensorRT-LLM, container, partner
Describes the serving stack for GPU-backed routes
policy.rules
priority list
Applies workspace, route, app, or customer rules before provider selection
policy.action
block, mask, warn, reroute, review
Controls the outcome when a policy or budget rule matches
usage.filter
time range, timezone, granularity, key, model, route
Keeps billing, logs, and capacity reports aligned
metadata.customer_id
string
Connects usage to customer, app, account, or procurement labels
metadata.trace_level
none, summary, full
Controls how much route evidence is returned or exported
metadata.export
logs, billing, security, partner
Routes approved metadata to the right operating system
playground.history
save, redact, export
Captures approved prompt tests with route and credit traces
ranking_scope
model, route, app, gpu_ready
Exports the comparison table behind a route decision
app_id
string
Attributes usage to a marketplace or private-catalog app
app.launch_profile
coding, research, media
Selects the evaluation and route evidence expected for an app
app.catalog_scope
public, private
Adds the appropriate listing or workspace-policy review gate
app.workload_shape
bursty, sustained
Maps the preview to serverless, reserved, or dedicated capacity
budget_key
string
Maps usage to a credit budget
budget.interval
daily, weekly, monthly, lifetime
Controls which workspace ceiling is checked before execution
batch.completion_target
flexible, standard, urgent
Maps offline work to a simulated serverless, burst, or reserved capacity preference
batch.endpoint
chat/completions, responses, messages, embeddings
Preserves the application's request contract while the queue plan is evaluated
batch.credit_ceiling
AI Token amount
Holds a batch plan for review when its simulated estimate exceeds the workspace ceiling
batch.result_mode
per_request
Keeps row-level outcomes inspectable so a held item does not erase eligible results
batch.return_plan_receipt
true, false
Returns request volume, estimate, budget gate, capacity lane, and lifecycle without submitting work
reconciliation.return_preview
true, false
Returns delivery state, attempts, provisional lines, and a planning receipt without changing billing policy
reconciliation.usable_output
observed, not observed
Separates delivery evidence from provider work before ledger review
observability.destination
webhook, warehouse, trace sink
Sends approved route events to a managed destination
return_trace
true, false
Returns route decision metadata
Route objectives
A useful router should expose simple objective presets, then translate them into provider order, service tier, policy checks, endpoint health, and credit ceilings.
Objective
Optimizes for
Best use
cost
lowest healthy credit exposure
support bots, batch jobs, broad experimentation
quality
highest route score and model fit
research, coding agents, executive answers
latency
fastest healthy provider and endpoint
interactive chat and customer-facing copilots
balance
weighted cost, quality, and latency
default production traffic and app backends
policy
retention, region, provider, and app rules first
enterprise, regulated, and private-catalog traffic
capacity
endpoint health and GPU lane availability first
serverless bursts, reserved lanes, and dedicated endpoints
Session continuity
The proposed Aurona preview reuses a compatible session route only while task fit, endpoint health, workspace policy, AI Token budget, and GPU capacity preference remain eligible. Every result is planning-only and does not pin production traffic.
Decision layer
Evidence
Planning use
Session context
session id, prior task, prior route
carry useful continuity evidence without treating the prior route as permanent
Task check
same workload or changed workload
reuse only when the next turn still fits the route
Health check
healthy or degraded
switch before a failed or saturated route breaks the session
Policy gate
unchanged or review required
hold before execution when workspace rules need another decision
AI Token gate
open or ceiling reached
stop the next turn before new credits are consumed
Capacity evidence
serverless or reserved GPU review
preview the lane transition without reserving capacity
Session continuity preview
const continuity = await fetch(
"https://api.aurona.ai/v1/routes/session-continuity-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
session_id: "agent-review-42",
prior: { task: "code", route: "aurona/code-balanced" },
next: { task: "analysis" },
evidence: {
route_health: "healthy",
policy_state: "unchanged",
budget_state: "open",
capacity_preference: "serverless"
},
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning-only contract. It returns reuse, switch, or hold evidence.
// It does not pin production traffic, debit AI Tokens, or allocate GPU capacity.Model shortlists
Aurona shortlist planning combines modeled catalog evidence with policy, lifecycle, economics, performance, and GPU readiness. The result is a simulated route draft, not a production change or binding price.
Decision layer
Signals
Planning use
Objective
balanced, quality, cost
Select the decision lens before candidates are scored
Evidence
quality, intelligence, design fit, usage
Keep benchmarks and market demand visible without treating either as absolute truth
Economics
input, output, request, media, cache
Estimate the full AI Token exposure for the intended workload
Performance
latency and throughput bands
Separate interactive, agent, batch, and media routes
Hard requirements
retention, region, parameter, lifecycle, capacity
Exclude incompatible supply before weighted scoring
Route draft
alias, fallback, budget, service tier
Turn the winning shortlist row into a simulated production plan
Workload evaluations
The proposed Aurona evaluation contract pins candidate routes to the same workload cases, checks tool behavior and policy fit, estimates AI Token exposure, and previews the capacity path without changing production traffic.
Evaluation layer
Evidence
Promotion use
Suite
saved prompts, expected behavior, tool assertions
measure the application workload instead of a generic benchmark
Pinned run
route, model version, harness, parameters
keep comparisons attributable to the candidate being tested
Release gates
quality floor, case pass rate, tool accuracy
block regressions before a route alias changes
AI Token evidence
suite estimate, candidate total, ceiling
compare quality and operating cost on the same workload
Policy evidence
workspace, region, provider, retention
exclude routes that cannot satisfy the launch contract
Capacity handoff
serverless, reserved, dedicated
promote sustained eligible traffic into a GPU-backed lane
Evaluation run preview
const evaluationRun = await fetch(
"https://api.aurona.ai/v1/evaluations/support-agent-v3/runs",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
candidates: ["aurona/agent-balanced", "aurona/auto-fast", "aurona/private-gpu"],
pin: { harness: "support-agent-v3", parameters: true },
release_gate: {
minimum_quality: 84,
require_tool_assertions: true,
maximum_ai_tokens: 40,
require_workspace_policy: true
},
traffic_forecast: { monthly_requests: 1500000 },
mode: "preview"
})
}
).then((response) => response.json());
// Preview-only contract: returns case evidence, failed gates, AI Token estimate,
// and a serverless, reserved, or dedicated capacity recommendation.Router metadata
Aurona route metadata should help developers debug provider choice while giving finance, security, app, and capacity owners the fields they need for exports.
Metadata
Fields
Used by
provider.selected
provider id, route id, service tier
developer trace and support review
session.continuity
reuse, switch, hold, prior route, selected route
multi-turn route review and agent debugging
fallback.reason
timeout, rate limit, policy, price, capacity
operations and route tuning
cache.trace
affinity, key hash, read, write, miss reason
cost review and cache-aware route tuning
policy.action
matched rule, scope, action, owner
security and enterprise approval
credit.trace
input, output, Fusion, app, endpoint, coupon
billing and finance export
endpoint.health
queue depth, warm cache, quota, saturation
serverless and dedicated routing
endpoint.promotion
route evidence, policy review, budget window, owner
simulated serverless-to-capacity handoff
app.attribution
app id, install id, customer id, meter
marketplace and private catalog settlement
metadata.labels
environment, region, procurement, project
search, reporting, and account operations
workload.label
dimension, value, label-set id, sampled state
Activity, Logs, finance, and capacity review
Tool route evidence
The proposed Aurona preview filters endpoints by tool-schema and workspace capacity requirements, then ranks simulated validity, workload evaluation, throughput, and AI Token evidence with a visible fallback.
Stage
Evidence
Planning use
Capability gate
tool schema, parallel execution, supported controls
exclude incompatible endpoints before scoring
Workspace gate
approved pool and capacity scope
keep public, reserved, and dedicated supply inside policy
Reliability evidence
schema-valid request rate and workload evaluation confidence
compare tool behavior instead of model popularity alone
Operating evidence
throughput index, AI Token estimate, capacity class
choose a balanced, reliability, or efficiency objective
Fallback evidence
second eligible endpoint and ordered review stages
make recovery reviewable before traffic runs
Planning receipt
simulated result, selected endpoint, fallback, evidence, and scope
save a decision without changing traffic, billing, or capacity
Tool route evidence preview
const toolRoutePreview = await fetch(
"https://api.aurona.ai/v1/routes/tool-evidence-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
route: "aurona/agent-balanced",
workload: "support-actions",
tools: { schema_contract: "strict", parallel: true },
objective: "reliability",
capacity_scope: "workspace-approved",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed preview contract. Returns simulated endpoint evidence and a fallback.
// It does not move traffic, debit credits, change billing, or allocate GPU capacity.Policy rules
Aurona can make enterprise controls operational by evaluating route, budget, region, app, prompt, and output rules before a request reaches a model.
Rule type
Match surface
Actions
Owner
Workspace rule
member, key, project, app
block, mask, warn, reroute
team governance
Route rule
model, provider, service tier
pin, fail closed, downgrade, escalate
routing control
Budget rule
customer, app, route, key
warn, block, require approval
finance control
Data rule
prompt, response, attachment
mask, redact, route privately
privacy review
Region rule
workspace, customer, endpoint
allow, deny, reroute
residency fit
App rule
install, runtime, meter
hold, approve, export
marketplace operations
Route eligibility
The proposed Aurona preview intersects workspace, member, and API-key scope, then applies AI Token budget, retention preference, and GPU capacity readiness. More restrictive scopes win; an empty pool or reached ceiling returns a reviewable result without selecting a route.
Decision layer
Inputs
Preview result
Workspace scope
approved providers, models, regions, and capacity
establish the broadest available pool
Member scope
role, team, project, and app policy
narrow the workspace pool for one operator
API-key scope
route, provider, budget, and environment
apply the most specific request credential limits
AI Token budget
used, ceiling, interval, and hard-stop state
block selection before provider or GPU assignment
Retention preference
standard or zero-retention planning attribute
remove incompatible simulated routes
Region execution path
gateway processing, model inference, and tool execution evidence
fail closed when any stage lacks the reviewed region
Capacity preference
shared, serverless, reserved, or dedicated
return only operationally suitable lanes
Decision receipt
effective pool, selection, fallback, and ordered stages
give developers and reviewers the same evidence
Eligibility preview
const eligibility = await fetch(
"https://api.aurona.ai/v1/routes/eligibility-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
workspace: "enterprise-prod",
member: "support-platform",
api_key: "support-prod-route",
route: "aurona/auto",
retention_preference: "zero_retention_preferred",
execution_region: "eu_review",
require_region_evidence: ["gateway", "inference", "tools"],
capacity_preference: "dedicated_or_serverless",
budget: { used: 58, ceiling: 80, interval: "month" },
mode: "preview"
})
}
).then((response) => response.json());
// Proposed preview response: effective providers, eligible Aurona routes,
// selected route, fallback, regional execution evidence, capacity class,
// budget result, and ordered stages. This is not a residency guarantee.
// An empty pool returns policy_review; a reached ceiling returns budget_blocked.Workspace change preview
The proposed Aurona preview resolves a workspace change against current keys, private apps, route aliases, budgets, and GPU capacity paths. It returns affected object identities and reasons while keeping every production control unchanged.
Stage
Evidence
Read-only result
Proposed scope
providers, budget window, capacity preference
A read-only change set
Dependency scan
API keys, private apps, route aliases, GPU paths
Stable object identities and current bindings
Effective diff
unchanged, budget recheck, provider removed, capacity mismatch
Reasons per affected object
Decision boundary
safe to review or hold for owner review
No assignment, credential, billing, or capacity mutation
Service tiers
A modern router should let teams choose the capacity class that matches the risk of the workload, from shared model access through dedicated GPU-backed routes.
Tier
Capacity path
Best for
shared
standard public model pool
experiments, prototypes, low-risk app traffic
priority
preferred provider order and health gates
production apps that need steadier latency
serverless
managed endpoint pool with autoscaling
variable workloads, app bursts, and endpoint trials
reserved
reserved throughput or GPU lane
high-volume routes with budget forecasts
dedicated
private model, tenant, or regional lane
regulated, private, or latency-sensitive workloads
sovereign
approved regional providers and private capacity
public sector, regulated, or residency-bound workloads
Endpoint lifecycle
The same route can move from playground validation into serverless inference, reserved throughput, dedicated endpoints, or partner GPU lanes as volume and risk change.
Stage
Use
Aurona control
Playground
prompt tests and parameter exploration
save approved history with cost and route trace
Serverless
managed pool for variable traffic
watch queue depth, warm cache, quota, and model health
Reserved
predictable route or app volume
attach budget forecast and throughput target
Dedicated
tenant or private endpoint
pin model, region, network, and service owner
Sovereign
regulated regional route
require approved vendor, region, and audit export
Partner lane
qualified GPU supplier capacity
settle usage through capacity and credit ledgers
Usage settlement preview
The proposed Aurona contract preserves each native meter, models an AI Token reservation, and reconciles only after complete provider-reported final usage arrives. Missing evidence leaves the wallet unchanged. All values are simulated and non-binding.
Evidence
Returned fields
Planning use
Native meter
input, output, cached, reasoning, request, image, megapixel, second
Preserve the provider-reported unit and billable identity
Reservation
request estimate + modeled translation weight + service tier
Hold a visible AI Token planning amount before execution
Final evidence
complete provider-reported usage lines
Fail closed when any required native meter is missing
Reconciliation
reservation compared with final measured usage
Explain the difference without silently rewriting the source meter
Capacity context
shared, priority, batch, or dedicated review
Carry economics toward GPU planning without allocating capacity
Planning receipt
profile, tier, evidence state, native line values
Make repeated read-only previews deterministic and reviewable
Settlement plan
const settlement = await fetch(
"https://api.aurona.ai/v1/credits/settlement-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
workload: "image",
service_tier: "priority",
usage_state: "provider_final",
native_meters: [
{ billable: "input_reference", unit: "image", estimate: 2, final: 2 },
{ billable: "output_image", unit: "image", estimate: 8, final: 8 },
{ billable: "output_pixels", unit: "megapixel", estimate: 16, final: 18 }
],
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning response: native meter lines, modeled AI Token reservation,
// final-usage evidence state, reconciliation preview, capacity context, and receipt.
// It does not debit a wallet, publish a rate card, or reserve GPU capacity.Credit headroom preview
The proposed Aurona contract reviews available AI Tokens, a workspace reserve, request exposure, current in-flight demand, and modeled route capacity before returning admit, queue, or review. All values are simulated and non-binding.
Evidence
Returned fields
Planning use
Credit evidence
available balance, workspace reserve, request estimate
hold before admission when spendable headroom is incomplete
Demand evidence
in-flight and requested concurrency
separate request pressure from the credit balance
Capacity evidence
shared, priority, or reserved-review lane
return a modeled ceiling without promising production throughput
Outcome
admit, queue, credit review, or evidence review
make the next operational step explicit and fail closed
Planning receipt
inputs, outcome, capacity class, simulated state
save a deterministic review record without changing traffic or billing
Admission plan
const headroom = await fetch(
"https://api.aurona.ai/v1/credits/headroom-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
available_ai_tokens: 120,
workspace_reserve: 20,
in_flight: 8,
requested_concurrency: 6,
estimated_ai_tokens_per_request: 2,
capacity_class: "shared",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning response: spendable credit headroom, modeled lane headroom,
// admit/queue/review outcome, ordered evidence, and deterministic receipt.
// It does not debit credits, set a production limit, or allocate capacity.Credit ledger
Credits become trustworthy when developers can see which tokens, media units, judge passes, route fees, and GPU lanes created the total.
Ledger field
Measured from
Settles
model.input_tokens
prompt and context tokens
provider model cost
model.output_tokens
completion tokens
provider model cost
fusion.judge_tokens
comparison and scoring tokens
Fusion orchestration
media.units
image, audio, video outputs
multimodal billing
gpu.capacity
serverless, reserved, or dedicated lane usage
compute settlement
credit.coupon
starter, promo, or enterprise credit adjustment
account balance
account.order
top-up, invoice, or committed-credit workflow
account operations
aurona.fee
routing, policy, trace, settlement
platform revenue
Credit accounts
A useful AI Token account view should keep promotional credits, top-ups, commitments, app settlement, endpoint usage, and partner capacity visible as distinct ledger lines.
Account line
What it represents
Where it applies
Paid balance
customer-funded AI Token credits
model tokens, Fusion, apps, endpoints, and capacity
Starter credit
trial or onboarding allocation
developer evaluation without publishing a binding rate card
Coupon credit
promo or enterprise adjustment
separate ledger line with source, owner, and expiration state
Top-up order
self-serve or invoice request
status, amount, workspace, tax, and payment path
Committed credits
contracted drawdown plan
budget forecasts, invoices, and procurement review
Partner settlement
compute or app supplier line
route, app, endpoint, region, and approval packet
Budget intervals
Aurona can make AI Token governance concrete by checking interval budgets, route ceilings, and app or customer caps before model, Fusion, or GPU-backed requests run.
Budget
Reset or scope
Why it matters
Daily
resets each workspace day
stop runaway spend from a single broken job or agent loop
Weekly
resets each operating week
smooth campaigns, app launches, evals, and team sprints
Monthly
resets each billing month
align credit drawdown with finance review and invoices
Lifetime
does not reset
hard cap trials, pilots, grants, and procurement-approved experiments
Route ceiling
checked per request
block expensive panels, premium models, or GPU lanes before execution
App/customer cap
checked per attribution key
keep marketplace installs and customer workspaces inside budget
Batch jobs
Aurona's proposed batch contract preserves Chat Completions, Responses, Messages, or Embeddings as an explicit request shape, separates offline work from interactive traffic, previews its AI Token ceiling, and returns a GPU-backed queue preference plus row-level lifecycle evidence before submission. Individual holds remain inspectable without erasing eligible rows. The contract and values shown here are simulated; planning does not submit work or allocate capacity.
Control
Example
Why it matters
Workload
text requests or embeddings
keep offline work separate from interactive and media routes
Protocol
Chat Completions, Responses, Messages, Embeddings
preserve the application contract instead of flattening every batch into one shape
Planning gate
request count, completion target, AI Token ceiling
review exposure before a queue accepts work
Capacity preference
serverless batch, burst GPU, reserved GPU
match flexible or urgent work to an explicit lane
Lifecycle
validated, queued, running, reconciling, completed
make asynchronous progress and partial failures inspectable
Row outcomes
completed or held for review
keep eligible results moving while individual rows remain independently inspectable
Plan receipt
stable simulated identifier and decision fields
save evidence without submitting work or allocating capacity
Batch plan preview
const plan = await fetch(
"https://api.aurona.ai/v1/batches?plan_only=true",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
endpoint: "/v1/responses",
route: "aurona/text-balanced",
request_count: 2400,
completion_target: "standard",
credit_ceiling: 5,
capacity_preference: "burst_gpu",
result_mode: "per_request",
return_plan_receipt: true
})
}
).then((response) => response.json());
// Proposed planning contract. plan_only does not submit work or allocate capacity.Capacity promotion preview
Aurona's proposed preview compares one native workload track across a completed observation window, useful utilization, traffic shape, model-runtime fit, target-region evidence, and normalized AI Token exposure. Missing evidence fails closed. The preview does not reserve GPUs, alter traffic, promise availability, publish pricing, or change billing.
Evidence
Example
Planning use
Completed window
14 or 30 observed days
Avoid promoting a short-lived launch spike as sustained demand
Native workload
routed tokens, audio minutes, or completed media seconds
Keep unlike operating units separate before comparison
Traffic shape
steady, variable, or bursty
Keep irregular demand in a flexible lane when reservation would create idle exposure
Useful utilization
observed percent plus workload-specific review floor
Compare productive capacity instead of raw GPU ownership
Runtime and region
verified model stack plus available target-region supply
Fail closed before a capacity plan reaches commercial review
Exposure comparison
normalized serverless and reserved AI Token indices
Model direction without publishing a rate, discount, or invoice
Planning receipt
normalized evidence and review outcome
Keep repeated previews deterministic without reserving GPUs
Capacity promotion plan
const preview = await fetch(
"https://api.aurona.ai/v1/capacity/promotion-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
workload: "agent-fleet",
evidence_window: { completed_days: 30 },
native_measure: "routed_tokens",
traffic_shape: "steady",
useful_utilization_percent: 68,
runtime_evidence: "verified",
region_evidence: "available",
return_planning_receipt: true
})
}
).then((response) => response.json());
// Proposed planning contract. Indices are simulated and normalized.
// This does not reserve GPUs, change routes, publish rates, or debit AI Tokens.Workspaces
Aurona workspaces separate personal builders, company teams, enterprises, and compute partners while keeping one control model for keys, routes, budgets, logs, credits, and invoices.
Workspace
Controls
Best use
Personal
owner, project keys, starter credits
developer trials and prototypes
Company
members, roles, invoices, budgets
team production traffic and app launches
Enterprise
SSO, audit, ZDR policy, procurement
governed model access and private routes
Compute partner
capacity profile, DD evidence, settlement
GPU supply mapped into Aurona routes
Agent connector
Aurona can expose a scoped connector for model discovery, ranking checks, credit visibility, docs lookup, and safe route tests while keeping production API calls on the normal API.
Surface
Returns
Developer value
Catalog
models, context, modality, price class, route score, lifecycle
choose a model without leaving the coding tool
Rankings
usage, latency, spend, task mix, app, and GPU-ready tables
compare current routes before migration
Credit state
balance, coupons, budget interval, and route ceiling
avoid accidental spend during development
Safe test call
scoped prompt test with redaction and trace
validate a model route before production code changes
Docs lookup
quickstarts, examples, errors, and endpoint schemas
answer integration questions from the editor
Workspace scope
project, key, role, and approved surfaces
keep connector sessions bounded to the current workspace
Agent policies
The Aurona Agent Route Lab turns coordinator choice, worker choice, allowed tools, child tasks, AI Token exposure, and GPU-backed capacity preference into one simulated policy contract.
Control
Example
Purpose
Coordinator route
aurona/agent-balanced
plans work and reconciles worker results
Worker route
aurona/worker-value
handles bounded child tasks on an approved route
Allowed tools
catalog, docs, safe test, app search
limits which platform surfaces the run may call
Child-task limit
1, 3, 5, or 8
caps delegation depth before execution
Run credit ceiling
workspace-defined AI Token value
stops the preview when the run reaches its budget
Capacity preference
serverless, reserved, or dedicated
maps agent work to an approved inference lane
Agent route policy
const policy = await fetch(
"https://api.aurona.ai/v1/agent-policies",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
name: "aurona/agent-balanced",
coordinator_route: "aurona/agent-balanced",
worker_route: "aurona/worker-value",
allowed_tools: ["catalog.read", "docs.lookup", "route.safe_test"],
child_task_limit: 3,
run_credit_ceiling: "0.60",
capacity_preference: "serverless_first",
mode: "preview"
})
}
).then((response) => response.json());
// Preview response: simulated route steps, tool decisions, credit estimate,
// policy checks, and selected GPU-backed capacity class.Workload labels
Aurona Workload Labels are a proposed, simulated workspace contract for attaching structured reporting dimensions to usage. Teams can preview sampling and an AI Token ceiling, then inspect the resulting labels in Logs and Activity without changing route behavior.
Control
Example
Purpose
Label dimensions
department, task, complexity, application, capacity
define the workspace reporting vocabulary
Sampling rate
10%, 25%, 50%, or 100%
preview credit exposure before broader coverage
AI Token ceiling
workspace-defined daily limit
bound the simulated labeling budget
Activity rollup
spend, tokens, requests, routes, apps, GPU lanes
compare workload shape with platform usage
Log evidence
label-set id, dimension, value, sampled state
filter and review individual simulated requests
Workload label preview
const labelSet = await fetch(
"https://api.aurona.ai/v1/workload-labels",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
name: "agent-operations-view",
workspace_scope: "aurona-launch",
label_dimensions: {
department: ["engineering", "product", "support", "operations"],
task: ["agent support", "code review", "research", "media workflow"],
capacity: ["serverless", "reserved", "dedicated"]
},
sampling_rate: 0.25,
ai_token_ceiling: { amount: 12, interval: "day" },
mode: "preview"
})
}
).then((response) => response.json());
// Preview-only contract: simulated labels can appear in Logs and Activity
// alongside route, app, AI Token, and GPU-capacity attribution.Console map
As customers move from first API call to production routes, the workspace needs clear surfaces for models, endpoint lifecycle, workflow execution, app runtime review, and finance exports.
Surface
Shows
Primary job
Dashboard
recent activity, usage trend, alerts, shortcuts
daily owner view
Model hub
catalog filters, model cards, examples, parameters
model selection and migration
Playground
prompt tests, multimodal runs, saved history
developer iteration and review evidence
Endpoints
serverless, reserved, dedicated, health, owner
capacity lifecycle management
Storage
model files, artifacts, logs, and rebuild inputs
private routes and workflow outputs
Workflow studio
node pipelines, sessions, reusable templates
GPU-backed app and media workflows
App runtime
hosted, self-hosted, serverless, dedicated
marketplace and private-catalog approval
Billing
credits, coupons, budgets, exports, settlement
finance reconciliation
Connector permissions
The proposed Aurona agent connector keeps catalog, rankings, docs, and credit checks read-only, while prompt or media tests require an explicit metered action on a separate bounded session key.
Capability
Permission
Returned or changed
Catalog and docs
read-only
models, endpoints, lifecycle, capabilities, examples, and errors
Rankings and task mix
read-only
usage, benchmarks, latency, app demand, and workload categories
Credit state
read-only
balance, budget window, coupon state, and route ceiling
Prompt test
explicit metered action
one scoped inference call with estimate, response, provider, and trace
Media test
explicit metered action
one image, audio, or video test only after cost preview
Run feedback
write to owned run
attach a category and note to a generation in the current workspace
Session key
time- and budget-bounded
separate connector access from application production keys
Runtime review
Marketplace and private-catalog apps need more than install buttons. Aurona can show where an app runs, how it is metered, which routes it uses, and what evidence was reviewed.
Runtime
Control
Best for
Hosted app
Aurona-managed runtime and route meter
builder launch with simple install flow
Self-hosted app
external runtime with signed callbacks
teams that keep execution in their own account
Serverless endpoint
managed GPU-backed inference pool
bursty agents, evals, and launch experiments
Dedicated endpoint
tenant lane with pinned model and owner
enterprise or high-volume private catalog apps
Review packet
playground runs, policy actions, samples, meter
catalog approval and enterprise procurement
Version state
draft, reviewed, listed, private, deprecated
safe app lifecycle and rollback
Launch packet
route evidence, AI Token meter, policy gate, capacity sample
review before listing or reserving capacity
Proposed flow: call /v1/apps/readiness-preview while iterating, then save the reviewed result through /v1/apps/launch-packets. Saving a packet does not publish a listing or allocate GPU capacity.
App signal registration
The proposed Aurona preview gives an app a stable simulated identity, then shows which public or workspace surfaces can use its route, AI Token, model-mix, and GPU capacity evidence. Workspace-only plans suppress every public surface.
Evidence
Fields
Planning use
App identity
stable app id, display name, category
join route and usage evidence without using a production API key as the app identity
Visibility
public catalog or workspace only
fail closed on public discovery, rankings, and model-app views
Usage evidence
requests, AI Tokens, model mix, time window
preview app demand without creating a real billing or ranking claim
Route context
route alias, environment, workspace
connect application demand to the model API and Fusion layer
Capacity context
serverless, reserved, dedicated review
show how app demand maps to GPU-backed inference capacity
Planning receipt
surface states, ordered checks, boundaries
save a deterministic preview without publishing or reserving capacity
App signal preview
const signalPlan = await fetch(
"https://api.aurona.ai/v1/apps/signal-registration-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
app_profile: "research-copilot",
visibility: "workspace",
environment: "staging",
route: "aurona/fusion-research",
capacity: "serverless",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning contract. Returns simulated surface, usage, AI Token,
// route, and capacity evidence. It does not publish, bill, or reserve capacity.Workspace runtime preview
Aurona can join workspace file scope, sandbox continuity, outbound network policy, model route, AI Token budget, and GPU capacity in one planning receipt. Missing scope or budget evidence fails closed.
Boundary
Evidence
Preview behavior
Workspace scope
approved file ids and app workspace
hold before route selection when the scope is missing
File boundary
isolated copies, explicit promotion review
separate durable workspace documents from temporary runtime output
Continuity
ephemeral or session reuse
declare whether later turns can see prior runtime files
Network policy
off or approved domains
keep outbound access explicit and reviewable
Route and credits
balanced or quality route, modeled AI Token budget
plan model work before any debit
GPU capacity
serverless pool or reserved review
hold restricted-budget plans without allocating capacity
Workspace runtime plan
const runtimePlan = await fetch(
"https://api.aurona.ai/v1/apps/runtime-preview",
{
method: "POST",
headers: {
Authorization: `Bearer ${AURONA_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
workspace_file_scope: "approved",
continuity: "session",
network_policy: "allowlist",
route_objective: "balanced",
ai_token_budget: "standard",
capacity: "serverless",
mode: "preview"
})
}
).then((response) => response.json());
// Proposed planning-only contract. It does not upload files, start a sandbox,
// debit AI Tokens, promise retention, change network policy, or allocate GPUs.Observability
Developers and enterprise buyers need route IDs, provider decisions, token burn, fallback events, latency, and policy outcomes.
Route trace
{
"route_id": "rt_aurona_fusion_01",
"model": "aurona/fusion",
"provider": "approved-provider",
"service_tier": "priority",
"key_mode": "customer_vault",
"session_id": "agent-thread-42",
"fallback_used": false,
"policy": { "data_collection": "deny", "region": "us" },
"credits": {
"input": "0.041",
"output": "0.088",
"fusion": "0.055",
"total": "0.184"
}
}Broadcasts
Aurona can make observability a first-class API surface by copying approved route events into logs, webhooks, warehouses, app dashboards, and finance systems.
Broadcast subscription
{
"sink": "usage-warehouse",
"events": ["request.created", "provider.selected", "credit.debited"],
"delivery": "signed_webhook",
"filters": {
"workspace": "prod",
"service_tier": ["priority", "reserved", "sovereign"],
"exclude": { "api_key": ["dev-sandbox"] }
},
"redaction": "metadata_only"
}Usage logs
Usage logs are more valuable when they connect model choice, provider health, key mode, token burn, app attribution, and policy decisions in one exportable event stream.
Event
Fields
Audience
request.created
route, model, workspace, customer, app
developer debugging
provider.selected
provider, service tier, key mode
routing transparency
cache.affinity
scope, provider, read, write, miss reason
cost and latency review
fallback.used
previous provider, next provider, reason
operations
api_key.usage_window
key, budget, spend, requests, trend
workspace owners
endpoint.health_sample
queue depth, warm cache, quota, saturation
router and capacity owners
billing.export.created
invoice, app, credit, coupon, and partner lines
finance
route.metadata.returned
provider, fallback, policy, credit, endpoint, labels
developer support and account operations
log.filter.applied
include, exclude, request id, workspace
support and audit
credit.debited
input, output, Fusion, GPU, app fee
finance and billing
broadcast.delivered
sink, event count, retry state
SRE and analytics
guardrail.triggered
rule, category, remediation
security and product
policy.blocked
rule, region, retention, vendor
enterprise review
policy.actioned
rule, scope, action, evidence, owner
security and route approval
playground.history.saved
prompt, modality, route, cost, redaction
developer iteration
Usage reports
Aurona reports should let owners filter by time window, timezone, key, model, route, endpoint mode, and export target without losing credit-ledger attribution.
Filter
Values
Why it matters
Time range
hourly, daily, monthly
compare logs, invoices, and endpoint usage consistently
Timezone
workspace local or UTC
keep finance exports and support review aligned
API key
project, route, endpoint, policy
show who created cost and which budget applied
Model and route
exact model, route alias, Fusion panel
separate raw provider cost from route value
Endpoint mode
serverless, reserved, dedicated
split variable inference from capacity commitments
Export target
billing, security, product, partner
send the right rows to each operating team
Creator and owner
member, service account, app, partner
resolve who created traffic or capacity cost
Credit source
paid, coupon, committed, partner
keep balances and adjustments auditable
Activity investigation preview
The proposed Aurona contract keeps the metric, dimensions, comparison window, saved-view scope, and exact request filters together. It joins AI Token, app, route, workspace, and GPU-lane evidence without exposing prompt content or changing live state.
Layer
Evidence
Review purpose
Metric contract
spend, requests, tokens, cache, latency, throughput
declare the measure before comparing rows
Dimension contract
workspace, app, route, model, key, workload, credit, capacity
join commercial and infrastructure evidence without flattening it
Comparison
current window plus adjacent completed window
separate a change signal from absolute usage
Request handoff
exact workspace, route, key, app, workload, credit, and lane filters
move from an aggregate anomaly to its simulated source requests
Saved-view scope
private preview or workspace preview
show intended visibility without persisting a dashboard
Planning boundary
no prompt content, no saved state, no ledger or capacity mutation
keep the website demonstration read-only
Activity investigation
const preview = await fetch(
"https://api.aurona.ai/v1/activity/investigation-preview",
{
method: "POST",
headers: {
Authorization: "Bearer <management-key>",
"Content-Type": "application/json"
},
body: JSON.stringify({
workspace_id: "workspace-aurona-launch",
metric: "ai_token_spend",
dimensions: ["app", "route"],
rollup: "day",
compare_previous_window: true,
selected_evidence_id: "req-aur-2048",
saved_view_visibility: "workspace"
})
}
);
// Proposed read-only contract. Returns simulated aggregate rows, comparison,
// exact request-log filters, saved-view scope, ordered stages, and a receipt.
// It does not save a view, reveal prompt content, debit credits, or reserve GPUs.Billing exports
Aurona billing exports should make it possible to reconcile usage across OpenAI-compatible model calls, Fusion runs, app meters, endpoint usage, coupons, and partner capacity.
Export
Fields
Audience
Credit ledger
paid, coupon, committed, adjustment
finance reconciliation
Model usage
fresh input, cache read, cache write, output, retry, provider
developer and cost review
Fusion usage
panel model, judge, final model, trace
route economics
App settlement
install, app, customer, builder, meter
marketplace operations
Endpoint usage
serverless, reserved, dedicated, region
capacity planning
Partner capacity
supplier, lane, health, invoice reference
compute settlement
Playground history
A saved playground run can carry model parameters, modality, route decision, credit estimate, redaction mode, and promotion status into docs, logs, or endpoint planning.
Stage
Use
Aurona control
Playground
prompt tests and parameter exploration
save approved history with cost and route trace
Serverless
managed pool for variable traffic
watch queue depth, warm cache, quota, and model health
Reserved
predictable route or app volume
attach budget forecast and throughput target
Dedicated
tenant or private endpoint
pin model, region, network, and service owner
Webhooks
Aurona should notify teams when budgets, providers, apps, invoices, or policy decisions need attention.
Event
When it fires
Audience
media.job.completed
asynchronous image, video, or audio output is ready
apps, workflow studio
media.job.failed
generation stops after validation, queue, or runtime
developers, operations
credit.threshold
budget reaches warning or hard limit
finance, product
route.fallback
request switched provider or route
operations
route.broadcast
trace is copied to an approved sink
operations, finance
route.price_alert
source model price crosses route ceiling
finance, product
provider.health
latency, errors, or capacity changed
SRE
app.usage
customer or app credit event
marketplace
invoice.created
billing period closes
finance
policy.blocked
request denied by enterprise rule
security
SDKs
OpenAI-compatible calls make migration easy. Native SDKs can expose Aurona-specific route, credit, app, and observability helpers.
SDK
Best for
Status
TypeScript
server apps, edge apps, agents
first-class
Python
data workflows, agents, notebooks
first-class
REST
any backend or workflow tool
stable
OpenAI SDK
swap base URL and API key
compatible
CLI
bootstrap keys, routes, and local profiles
planned
MCP
agent access to model catalog, docs, and test routes
planned
Webhooks
billing and operations events
signed
Errors
Clear errors make Aurona feel production-grade: credits, policy, provider capacity, and route configuration should fail with useful guidance.
HTTP
Code
Developer action
400
invalid_request
Fix malformed route, model, or parameter
401
unauthorized
Check API key and environment
402
credit_required
Add credits or increase budget
403
policy_denied
Request violates provider, region, or retention policy
429
rate_limited
Use fallback, reserved lane, or retry schedule
503
provider_unavailable
Fallback was unavailable or disabled
Guides
Aurona needs clear docs across access, credits, routing, policy, compute, apps, and observability so developers can move from first call to production with confidence.
Swap your base URL, add an Aurona API key, choose aurona/auto, and send your first request.
Expose live model catalog, rankings, docs, route tests, and local project setup to coding agents and terminal workflows.
Give coding assistants scoped access to catalog, rankings, credits, docs, and safe test routes without exposing production secrets.
Preview coordinator and worker routes, tool allowlists, task limits, credit ceilings, and GPU capacity preference as one governed contract.
Point Codex, Claude Code, Cursor, OpenClaw, or internal agent runners at Aurona routes with scoped keys and budgets.
Browse copy-ready examples for text, image, video, audio, vision-language, embeddings, and 3D model routes.
Pull usage, spend, latency, benchmark, app, and GPU-readiness tables into planning reviews and route migration notes.
Run a panel of models, compare outputs with a judge, synthesize one answer, and return a route trace.
Sort by price, throughput, or latency while controlling fallbacks, provider order, data policy, and parameters.
Define priority rules that block, mask, warn, reroute, or review requests before provider selection.
Track usage by app, customer, environment, route, model, provider, and GPU lane.
Set prepaid credits, committed spend, alerts, hard stops, customer budgets, and route-level cost ceilings.
Explain daily, weekly, monthly, lifetime, route, app, and customer budget checks before requests execute.
Install public or private apps with route aliases, customer budgets, app usage logs, and builder settlement events.
Preview model evidence, route policy, an AI Token meter, and serverless or reserved capacity before a listing is published.
Estimate model, Fusion, app, fallback, and GPU lane exposure before approving production traffic.
Attach reserved inference lanes and advanced GPU supply for workloads that need predictable throughput.
Register serverless pools, dedicated endpoints, queue policy, warm-cache state, quota, and runtime metadata.
Show the dashboard, model hub, playground, endpoints, storage, workflow, runtime, and billing surfaces a workspace owner expects.
Save prompt tests, multimodal runs, model parameters, route decisions, and cost traces for later review.
Separate model tokens, Fusion, coupons, app settlement, endpoint usage, and partner capacity into finance-ready rows.
Add prompt, output, sensitive-data, vendor, and region checks as routing inputs before a provider sees traffic.
Configure retention, region, approved vendors, audit logs, SSO, spend controls, and fail-closed fallback.
Inspect route decisions, token burn, latency, provider health, fallback behavior, and customer usage.
Subscribe to credit thresholds, route failover, provider health, budget events, invoice activity, and policy blocks.
Advanced features
The docs go beyond the first API call into routing, fallbacks, tools, ZDR, attribution, service tiers, AI Token controls, and GPU capacity.
Feature
API surface
Why developers need it
Protocol route preview
/v1/routes/protocol-preview
Prefer native request formats, expose reviewed adapters, and verify an ordered fallback chain under one shared request contract
Request reconciliation preview
/v1/usage/reconciliation-preview
Review usable output, provider attempts, approved fallback, auxiliary work, and provisional AI Token lines without issuing credits or changing billing policy
Media job planning
/v1/media/jobs
Validate output controls, preview AI Token exposure, choose a GPU lane, and return synchronous output or an asynchronous job receipt
Media capability discovery
/v1/media/models
Discover output types and endpoint-specific resolution, aspect ratio, reference input, format, streaming, and delivery support
Model fallbacks
fallback: auto or approved-only
Switch providers or routes when capacity, price, or policy changes
Cache-aware routing
/v1/routes/:id/cache-policy
Set cache affinity and scope, then inspect provider-reported cache evidence without bypassing route policy
Provider key vault
provider.key_mode: customer_vault
Let customers attach approved provider credentials while Aurona keeps route, policy, and usage records
Management API
/v1/management/keys
Automate project, route, budget, and usage-log administration
API key detail
/v1/api-keys/:id/usage
Show per-key charts, budget progress, log shortcut, rotation state, and workspace ownership
Router metadata
/v1/routes/:id/metadata
Return provider, fallback, policy, credit, endpoint, app, and customer evidence when requested
Policy rule engine
/v1/policy/rules
Define priority-based block, mask, warn, reroute, and review actions for production traffic
Policy action log
/v1/policy/actions
Export matched rules with scope, owner, request id, evidence, and downstream billing impact
Observability destinations
/v1/observability/destinations
Manage Datadog, Langfuse-style, warehouse, webhook, and billing-export sinks from one surface
MCP model bridge
/v1/mcp
Expose live catalog, rankings, docs, and safe test inference to coding agents
Agent connector
/v1/agent-connectors
Create scoped sessions for coding tools to read catalog, rankings, docs, and credit state
Agent route policies
/v1/agent-policies
Bound coordinator and worker routes, tools, child tasks, AI Token exposure, and capacity preference
Workload labels
/v1/workload-labels
Preview workspace-scoped dimensions, sampling, AI Token ceilings, log tags, and Activity rollups
Model capability discovery
/v1/models/capabilities
Filter by modality, parameters, region, age, provider-declared training metadata, and endpoint class
Catalog query preview
/v1/models/query-preview
Filter hard capabilities, place missing evidence last, resolve a canonical route, and return a planning receipt
Ranking feed preview
/v1/datasets/rankings-preview
Return versioned ranking rows with explicit measure, window, opt-in scope, portable citation, AI Token, GPU capacity, and trend-comparison evidence
Model shortlist planning
/v1/model-shortlists
Rank eligible routes by balanced, quality, or cost objectives after hard policy and capacity exclusions
Workload evaluation
/v1/evaluations/:id/runs
Compare pinned route candidates on application prompts, tool assertions, quality, AI Token use, policy, and capacity evidence
Model lifecycle
/v1/models/lifecycle
Expose preview state, version pinning, retirement date, replacement route, and workspace owner
CLI bootstrap
aurona login + aurona route init
Help developers configure keys, route aliases, budgets, and local profiles quickly
Coding tool setup
OpenAI-compatible endpoint config
Provide copy-ready setup for Codex, Claude Code, Cursor, OpenClaw, and internal agent runners
Model quickstarts
/v1/model-quickstarts
Group text, image, video, audio, and 3D examples by model capability
Tool calling
tools + require_parameters
Route only to models and providers that support function calling
Structured outputs
response_format + schema
Keep JSON, scoring, and app automations reliable
ZDR
data_policy: zero_retention
Restrict prompts to providers and lanes with compatible retention rules
App attribution
app_id + customer_id
Settle usage to app builders, enterprise catalogs, and customer budgets
Service tiers
service_tier + compute_lane
Choose public, priority, reserved, or dedicated capacity
Route objectives
/v1/route-objectives
Reuse cost, quality, latency, balance, policy, and capacity presets across apps
Budget intervals
/v1/workspaces/:id/budgets
Enforce daily, weekly, monthly, lifetime, route, app, and customer ceilings before execution
Cost simulator
/v1/cost-simulations
Preview model, Fusion, app, and GPU credit exposure before routing production traffic
Guardrails
/v1/guardrails
Block prompt injection, sensitive data, unsafe outputs, or disallowed vendors before provider selection
Observability broadcast
/v1/observability/broadcasts
Send route traces into logs, webhooks, finance systems, and monitoring tools
Usage logs
/v1/observability/events
Inspect token burn, app attribution, fallback, policy, provider-health events, and request-id filters
Rankings export
/v1/rankings
Fetch usage, spend, latency, benchmark, app, and GPU-readiness tables for internal review
App marketplace install
/v1/apps/installs
Attach route aliases, app meters, customer budgets, and usage-log shortcuts during install
App runtime review
/v1/apps/reviews
Attach saved playground runs, runtime metadata, policy actions, and meter evidence before listing
Workspace menu state
/v1/workspaces/:id/menu
Show credits, labs, keys, logs, routes, and billing shortcuts in the account surface
Console surface map
/v1/console/surfaces
Expose dashboard, model hub, playground, endpoints, storage, workflow, runtime, and billing shortcuts
GPU endpoint tester
/v1/gpu-lanes/:id/test
Preview queue depth, model cache, latency, and capacity class before routing production traffic
Inference endpoints
/v1/inference/endpoints
Register serverless pools, dedicated endpoints, model cache, quota, health, and runtime metadata
Endpoint lifecycle
/v1/inference/endpoints/:id/lifecycle
Move a route from playground testing to serverless pool, reserved lane, or dedicated endpoint
Billing exports
/v1/billing/exports
Separate model tokens, Fusion, coupons, app settlement, endpoint usage, and partner capacity lines
Agent runtime path
/v1/apps/runtimes
Choose hosted, self-hosted, serverless GPU, or dedicated endpoint before an app listing is reviewed
Filter grammar
logs.filter
Support include and exclude chips for model, provider, key, workspace, route, and event type
Router metadata
metadata object
Attach customer, environment, region, and procurement labels to every request
Aurona.ai
API
OpenAI-compatible
Billing
AI Token credits
Capacity
GPU-backed routes