Aurona.ai

Developer documentation

Build on the AI compute network.

Start with one OpenAI-compatible request, then add intelligent routing, AI Token controls, policy, and GPU capacity.

OpenAI-compatibleAI Token ledgerGPU route controls

First routed request

TypeScript · OpenAI-compatible

const response = await fetch("https://api.aurona.ai/v1/chat/completions", {
  method: "POST",
  headers: { Authorization: `Bearer ${AURONA_API_KEY}` },
  body: JSON.stringify({
    model: "aurona/auto",
    route: { objective: "balance" },
    messages: [{ role: "user", content: "Build an agent." }]
  })
});
200 · route readyaurona/auto · balance
Browse documentation

Developer signals

Docs should make the platform feel buildable immediately.

The strongest API docs show both the familiar starting point and the differentiated controls: routes, credits, policy, GPU lanes, traces, and events.

Compatibility

OpenAI SDK

Start by changing base URL and API key, then add route, budget, policy, and trace fields.

Route config

JSON object

Document provider sort, fallback, data policy, service tier, budget key, compute lane, and return trace.

Ledger events

credits

Expose model, Fusion, media, GPU, app, and platform fee events for billing transparency.

Operations

broadcast

Send route traces into webhooks, logs, observability pipelines, billing exports, and app dashboards.

Account console

keys + credits

Document workspace menus for scoped keys, credits, coupons, usage reports, endpoint state, and labs.

Agent connector

docs + live data

Let coding assistants inspect model catalog, route prices, credit state, rankings, and safe test calls through a scoped connector.

Agent policies

route + tools

Bound coordinator and worker routes, allowed tools, child tasks, credit exposure, and compute preference before a run.

Budget windows

daily to lifetime

Show how workspace budgets can enforce daily, weekly, monthly, and lifetime credit ceilings before requests run.

Integration paths

Choose the lightest path that fits the workload.

Start with a direct request, preserve an existing OpenAI SDK integration, or give a scoped agent live catalog and docs context. Every path converges on the same Aurona routes, AI Token ledger, workspace policy, and GPU-backed capacity model.

Path

Starting point

Aurona control

Best use

Direct REST

Any server runtime; no additional client dependency

Full request control for exact models, Aurona routes, streaming, tools, and response metadata

Teams building a custom API layer or testing the first request

OpenAI SDK

Existing OpenAI-compatible application

Swap the base URL and key, then add Aurona route, budget, policy, and trace fields when needed

Teams migrating an established chat, agent, or tool-calling integration

Agent connector

Scoped coding tool or internal agent session

Read catalog, rankings, docs, and credit state; keep metered tests explicit and bounded

Builders who need live platform context while they implement or review a route

Project bootstrap lab · simulated

Turn onboarding choices into one reviewable packet.

Plan the integration path, scoped access, masked environment manifest, AI Token ceiling, live build context, verification steps, and GPU-capacity handoff before any production action.

Application profile

Bootstrap preview

API application

OpenAI-compatible API · Personal sandbox

Within preview ceiling

Key ownership

Personal project key plan

Modeled use

6 / 12 AI Tokens

Capacity

Serverless first

Masked environment manifest

AURONA_API_BASE
https://api.aurona.ai/v1
AURONA_API_KEY
<create-scoped-key-after-review>
AURONA_WORKSPACE
personal-sandbox
AURONA_ROUTE
aurona/auto

Live build context

  • ReadCatalog and model lifecycle
  • ReadRankings and route evidence
  • ReadAI Token balance and budget state
  • ReadDocs and safe test calls

Verification and handoff

  1. 01Resolve the model and route catalog
  2. 02Send one redacted test request
  3. 03Inspect route and AI Token evidence

Start with a managed route and watch workload evidence

Simulation only: this packet does not create or rotate credentials, does not write environment files, does not enable billing, does not allocate GPU capacity, and does not change production traffic.

receipt · bootstrap-preview:api-app:personal-sandbox:serverless-first:12

Route onboarding

Test a familiar request, then add only the controls the workload needs.

This simulated planning path keeps early integration simple while making the handoff to route review, workspace controls, and GPU-backed capacity legible.

Step

What the builder does

What Aurona keeps visible

Request builder

Send one OpenAI-compatible test with an exact model or aurona/auto.

Save the prompt, parameters, and credit estimate as a simulated planning path.

Route evidence

Compare catalog fit, ranking signals, provider policy, fallback, and budget scope.

Keep the selected route explainable before it receives app or production traffic.

Capacity promotion

Move a validated workload through serverless, reserved, or dedicated endpoint choices.

Review modeled endpoint health, ownership, and budget context; no capacity commitment is created here.

Model discovery

Query one catalog across modalities and route requirements.

Filter model supply by output type and supported request controls, then sort by economics, context, performance, demand, or catalog age. Workspace aliases can follow an approved upgrade policy while production routes remain version-pinned.

Catalog query

GET /v1/models
Show code7 lines
const models = await fetch(
  "https://api.aurona.ai/v1/models?output_modalities=text,image&supported_parameters=tools&sort=throughput-high-to-low",
  { headers: { Authorization: `Bearer ${AURONA_API_KEY}` } }
).then((response) => response.json());

// Use a workspace alias for controlled upgrades; pin exact versions in production.
const model = "aurona/latest-balanced";

Catalog query preview

Make catalog discovery reproducible before it becomes routing policy.

The proposed Aurona preview intersects hard capability, capacity, provider-declared data use, and access requirements before sorting visible evidence or resolving a workspace alias. It returns excluded reasons and a stable planning-only receipt. It does not change aliases or policies, send traffic, debit AI Tokens, publish models, or allocate GPU capacity.

Evidence

Returned fields

Planning use

Hard filters

output modality, required parameter, approved capacity

Exclude incompatible supply before any sort runs

Provider-declared conditions

data-use class and access class

Fail closed or route supply into workspace review before alias resolution

Excluded evidence

canonical ID and review reason

Show why otherwise-capable supply was removed before scoring

Sort evidence

AI Token estimate, throughput, latency, or demand

Place unmeasured candidates last instead of treating missing data as zero

Canonical identity

workspace alias → immutable route version

Keep discovery convenient while production evidence remains attributable

Lifecycle

preview, current, maintenance, retiring

Show maturity and replacement context alongside capability data

Planning receipt

normalized query + eligible canonical IDs

Make the same read-only query reproducible across workspace reviews

Capacity path

serverless or dedicated review

Carry model demand toward GPU planning without allocating capacity

Catalog query plan

POST /v1/models/query-preview
Show code24 lines
const preview = await fetch(
  "https://api.aurona.ai/v1/models/query-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      output_modality: "text",
      required_parameter: "tools",
      sort: "throughput_high_to_low",
      alias_policy: "workspace_alias",
      capacity: "serverless",
      data_use_policy: "no_contribution_only",
      access_policy: "standard_only",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed read-only response: eligible count, canonical route IDs,
// provider-declared data use, access review, exclusions,
// AI Token estimate, capacity path, and planning-only receipt.

Variant contract preview

Keep delivery, reasoning, and metering explicit when a family has multiple routes.

The proposed Aurona preview resolves an immutable simulated variant only after its delivery mode, reasoning class, native meter, task contract, modeled AI Token amount, and capacity path are declared. A missing or incompatible field holds the alias instead of silently substituting a different execution contract.

Evidence

Returned fields

Planning use

Family identity

stable family plus distinct variant ID

Keep discovery convenient without flattening operationally different routes

Delivery mode

interactive request or deferred job

Fail closed before a synchronous client receives an asynchronous contract

Decision shape

yes/no, allowed choice, or bounded score with abstention policy

Keep typed decision workloads distinct from generated prose and verify schema fit before routing

Reasoning class

standard or deliberate review

Keep higher-reasoning supply and its meter explicit

Native meter

input, output, reasoning, character, image, second, or request

Preserve the source unit before modeled AI Token translation

Task contract

synchronous response or task ID plus polling

Make recovery and status handling visible before integration

Capacity path

serverless, shared batch, or reserved review

Join execution mode to GPU planning without allocating capacity

Planning receipt

normalized request and selected immutable variant

Make alias resolution reviewable and reproducible

Variant contract plan

POST /v1/models/variant-preview
Show code20 lines
const preview = await fetch(
  "https://api.aurona.ai/v1/models/variant-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      family: "aurona/atlas",
      delivery_mode: "deferred",
      reasoning_class: "standard",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning-only response: immutable variant ID, native meter,
// synchronous or asynchronous task contract, modeled AI Tokens,
// capacity path, exclusions, and deterministic receipt.

Protocol route preview

Preserve the application contract before selecting supply.

The proposed Aurona preview evaluates Chat Completions, Responses, Anthropic Messages, and asynchronous media jobs as explicit route inputs. Native support ranks first; reviewed adapters remain visible; approved fallback attempts share one request contract. Any feature, adapter, capacity, or fallback loss produces a hold.

Evidence

Returned fields

Planning use

Incoming contract

Chat Completions, Responses, Anthropic Messages, or asynchronous video job

Resolve the route from the protocol the application already uses

Hard feature

tools, structured output, reasoning, streaming, media references, or async result

Exclude any path that cannot preserve the required behavior

Translation state

native or reviewed adapter

Prefer native support and make approved conversion visible

Fail-closed hold

protocol, feature, adapter, or capacity mismatch

Never silently relax the application contract

Commercial evidence

modeled AI Token estimate

Compare compatible paths without publishing a rate card

Capacity path

serverless or dedicated review

Join protocol compatibility to GPU planning without allocating capacity

Fallback contract

disabled or approved chain; shared request fields across attempts

Reject per-attempt overrides that could silently change application behavior

Ordered attempts

primary, fallback-1, fallback-2, trigger reasons

Hold when no second route preserves every protocol, feature, policy, and capacity gate

Protocol compatibility plan

POST /v1/routes/protocol-preview
Show code23 lines
const preview = await fetch(
  "https://api.aurona.ai/v1/routes/protocol-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      protocol: "anthropic_messages",
      required_feature: "tools",
      adapter_policy: "reviewed_adapter",
      capacity: "all",
      fallback_policy: "approved_chain",
      attempt_policy: "shared_contract",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning response: native and reviewed-adapter candidates,
// ordered fallback roles, shared-contract evidence, exclusions,
// modeled AI Token evidence, capacity, and receipt.

Ranking feed preview

Carry ranking evidence into internal tools without losing its meaning.

The proposed Aurona feed returns simulated model, route, and app rows with a declared measure, freshness, version, schema, canonical endpoint, portable citation, visibility scope, AI Token context, and GPU capacity path. Public previews fail closed on workspace-only evidence, while provenance metadata makes internal exports attributable without asserting reuse rights.

Evidence

Returned fields

Planning use

Dataset identity

as-of timestamp, version, schema, format, canonical endpoint

Keep every exported ranking snapshot attributable and reproducible

Portable citation

publisher, title, version, measure, window, scope, simulated origin

Carry human-readable provenance into internal reports without implying reuse rights

Measure semantics

modeled AI Token volume or request-share index

Separate adoption evidence from quality, latency, and benchmark claims

Visibility scope

workspace view or public opt-in only

Exclude workspace-only rows before ranking a public preview

Surface filter

models, routes, apps, or all eligible

Return only the operating surface requested by the consumer

Ranking rows

absolute rank, stable id, modeled requests

Preserve ordering after every hard filter

Capacity evidence

serverless, reserved review, dedicated review

Join demand signals to GPU planning without allocating capacity

Ranking feed query

GET /v1/datasets/rankings-preview
Show code9 lines
const feed = await fetch(
  "https://api.aurona.ai/v1/datasets/rankings-preview?surface=routes&scope=workspace-only&metric=ai-token-volume&window=30d",
  { headers: { Authorization: `Bearer ${AURONA_API_KEY}` } }
).then((response) => response.json());

// Proposed planning response: ranked rows plus as_of, version, resolved window,
// dataset manifest, portable citation, measure semantics, visibility scope,
// capacity evidence, and a stable receipt.
// It does not publish traffic, change visibility, debit credits, or reserve GPUs.

Native meter ranking preview

Compare one measure without flattening every kind of AI work.

The proposed Aurona preview ranks simulated routes only when measure, native unit, completed window, and visibility scope match. Tokens, images, tool calls, media time, and GPU hours remain separate. A within-track index shows relative position without inventing a cross-unit score.

Evidence

Returned fields

Planning use

Track contract

measure, native unit, completed window

Keep unlike usage evidence out of the same ordering

Native units

tokens, images, tool calls, audio seconds, video seconds, GPU hours

Preserve how each product surface actually measures work

Within-track index

leader equals 100; every other row is relative to that leader

Compare position without inventing cross-unit conversion

Visibility gate

public opt-in or workspace review

Suppress private evidence before a public ranking is calculated

Freshness

latest completed evidence bucket in UTC

Distinguish data freshness from page-render time

Capacity evidence

serverless, reserved review, dedicated review

Connect demand to GPU planning without allocating capacity

Planning receipt

normalized track query and eligible stable IDs

Make the same read-only ordering reproducible

Native meter ranking plan

POST /v1/rankings/native-meter-preview
Show code21 lines
const preview = await fetch(
  "https://api.aurona.ai/v1/rankings/native-meter-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      measure: "tool-activity",
      native_unit: "tool-calls",
      window: "7d",
      visibility: "public",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning-only response: ranked rows in one native unit,
// a within-track index, completed-window freshness, visibility and
// capacity evidence, exclusions, and a deterministic receipt.

Multimodal evaluation preview

Route image workloads with visible capability evidence.

The proposed Aurona preview compares simulated image routes against the same prompt family and evaluation floor, then ranks only eligible supply by quality, generation speed, or modeled AI Token efficiency. Capacity mismatches fail closed before a route is selected.

Evidence

Returned fields

Planning use

Shared prompt family

exact text, editing, consistency, counting, spatial

Compare candidates against the same observable capability

Evidence floor

case count and evaluation version

Keep small or stale samples out of route selection

Objective

quality, generation speed, or AI Token efficiency

Declare the ranking goal instead of hiding a blended score

Native evidence

pass rate, seconds, modeled AI Tokens

Preserve each measure behind the selected route

Capacity scope

serverless, reserved review, dedicated review

Fail closed when the requested GPU path has no eligible route

Planning receipt

normalized gates and eligible route IDs

Make a read-only comparison repeatable before production review

Multimodal route plan

POST /v1/evaluations/multimodal-route-preview
Show code21 lines
const preview = await fetch(
  "https://api.aurona.ai/v1/evaluations/multimodal-route-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      modality: "image",
      prompt_family: "exact-text",
      objective: "quality",
      minimum_cases: 12,
      capacity: "reserved-review",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning response: eligible and excluded routes, native evidence,
// capacity scope, selected route preview, evaluation version, and receipt.

Model alias lifecycle

Keep upgrades current without making production movement invisible.

Aurona aliases separate discovery from release control. Moving and preview aliases reduce catalog churn, while pinned versions and review evidence keep production routes attributable across policy, credits, evaluations, and capacity.

State

Example

Control

Best use

Moving alias

aurona/latest-balanced

Follows a workspace-approved family policy

Exploration and controlled development environments

Preview alias

aurona/preview-balanced

Exposes the next candidate to saved evaluations, policy checks, and cost review

Pre-production comparison without changing the stable route

Pinned version

aurona/balanced-2026-08

Keeps model, provider requirements, parameters, and capacity class attributable

Production routes that need deliberate releases and repeatable evidence

Retiring alias

aurona/balanced-2026-05

Shows replacement route, owner, evaluation state, and retirement window

Planned migration with logs and fallback evidence

Review before promotion

preview → stable

Checks workload quality, tools, AI Token budget, workspace policy, and GPU lane fit

A simulated review that does not move production traffic

API reference

One API surface for access, routing, and settlement.

Aurona keeps the first request familiar while adding the controls AI companies need as they scale.

Advanced routed request

POST /v1/chat/completions
Show code37 lines
const response = await fetch(
  "https://api.aurona.ai/v1/chat/completions",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: "aurona/auto",
      route: {
        objective: "balance",
        optimize: ["price", "latency", "policy"],
        cost_quality_tradeoff: 6,
        fallback: "auto",
        session_id: "agent-thread-42",
        data_policy: "zero_retention",
        service_tier: "serverless",
        endpoint_type: "managed_pool",
        budget_key: "prod-agent",
        budget_interval: "monthly",
        policy_rules: ["mask-sensitive-metadata", "approved-provider-order"],
        ranking_scope: "route",
        app_id: "internal-agent-workbench"
      },
      usage_filter: {
        timezone: "workspace",
        granularity: "hourly"
      },
      playground_history: {
        save: true,
        redaction: "metadata_only"
      },
      messages: [{ role: "user", content: "Build an agent." }]
    })
  }
);

Endpoint

Use

Layer

Area

/v1/chat/completions

Chat, agents, tools, streaming

OpenAI-compatible

core

/v1/models

Model and route catalog

provider-aware

catalog

/v1/models?output_modalities=

Filter text, image, audio, and embedding outputs

catalog

discovery

/v1/models/query-preview

Preview capabilities, provider-declared supply conditions, canonical identity, credits, and capacity

catalog

planning

/v1/models/variant-preview

Preview delivery mode, reasoning class, native meter, task contract, credits, and capacity

catalog

planning

/v1/routes/protocol-preview

Preview native or reviewed-adapter protocol support plus a contract-safe fallback chain

routing

planning

/v1/models/capabilities

Capability, region, training-review, and parameter filters

catalog

discovery

/v1/models/lifecycle

Release state, version policy, retirement date, and replacement route

catalog

operations

/v1/fusion/runs

Panel, judge, synthesis

multi-model

quality

/v1/routes

Create and manage route aliases

routing

control

/v1/routes/decide

Simulate task classification, approved-pool scoring, fallback, and decision receipt

routing

planning

/v1/routes/session-continuity-preview

Preview session route reuse, switching, policy holds, budgets, and GPU lane evidence

routing

planning

/v1/routes/tool-evidence-preview

Rank tool-capable endpoints with schema, evaluation, throughput, credit, and capacity evidence

routing

planning

/v1/routes/eligibility-preview

Preview layered workspace, member, key, budget, regional execution, retention, and capacity eligibility

routing

governance

/v1/route-objectives

Reusable cost, quality, latency, balance, and policy objective presets

routing

control

/v1/routes/:id/metadata

Route evidence, provider choice, policy action, credit trace

routing

trace

/v1/management/keys

Organization and automation keys

admin

control

/v1/api-keys/:id/usage

Per-key usage, budgets, and spend charts

account

control

/v1/api-keys/:id/logs

Pre-filtered logs for one API key

account

trace

/v1/mcp

Live model catalog and route testing for agents

tool server

developer

/v1/agent-connectors

Scoped connector sessions for coding tools and internal agents

developer

setup

/v1/projects/bootstrap-preview

Simulate integration, scoped access, AI Token budget, and capacity handoff

developer

planning

/v1/projects/context-manifest

Return a masked environment and live-context manifest without issuing credentials

developer

planning

/v1/agent-policies

Coordinator, worker, tool, budget, and capacity controls

agent runtime

control

/v1/workload-labels

Workspace label dimensions, sampling, and usage rollups

observability

governance

/v1/cli/sessions

CLI login, profile, and route bootstrap

developer

setup

/v1/provider-keys

Customer-owned provider key vault

routing

privacy

/v1/provider-keys/route-preview

Preview masked credential scope, health, priority, fallback, and capacity

routing

planning

/v1/model-quickstarts

Copy-ready model examples by modality

docs

onboarding

/v1/model-shortlists

Compare evidence, policy, lifecycle, cost, latency, and GPU readiness

catalog

planning

/v1/evaluations

Create application-specific prompt and tool-assertion suites

evaluation

planning

/v1/evaluations/:id/runs

Run pinned route candidates against the same workload cases

evaluation

evidence

/v1/evaluation-gates

Preview quality, tool, policy, credit, and capacity promotion gates

routing

control

/v1/rankings

Model, route, app, and GPU-ready ranking feeds

market data

rankings

/v1/rankings/native-meter-preview

Rank one declared measure, native unit, window, and visibility scope at a time

market data

planning

/v1/datasets/rankings-preview

Preview versioned ranking rows, measure semantics, visibility scope, citation metadata, and capacity evidence

market data

planning

/v1/evaluations/multimodal-route-preview

Compare shared prompt-family evidence before previewing an image route

evaluation

planning

/v1/policy/rules

Priority rules for prompt, response, provider, route, and budget actions

enterprise

policy

/v1/policy/actions

Block, mask, warn, reroute, or hold requests for review

enterprise

policy

/v1/credits

Wallet, ledgers, balances

token settlement

billing

/v1/credits/settlement-preview

Translate native usage meters into an AI Token reservation and reconciliation preview

token settlement

planning

/v1/credits/headroom-preview

Preview request admission from AI Token reserve, in-flight demand, and route capacity evidence

token control

planning

/v1/budgets

Limits, alerts, hard stops

spend control

finance

/v1/workspaces/:id/budgets

Workspace budget intervals and enforcement state

spend control

finance

/v1/batches

Plan and submit asynchronous text or embedding work

inference

batch

/v1/batches/:id

Inspect validation, queue, execution, reconciliation, and result state

operations

batch

/v1/cost-simulations

Estimate route and model spend

planning

finance

/v1/usage/reconciliation-preview

Simulate delivery outcome, attempts, fallback, and provisional AI Token lines

planning

finance

/v1/usage/requests/:id

Inspect request outcome, route evidence, attempts, and reconciled usage state

operations

trace

/v1/activity/investigation-preview

Preview aggregate queries, prior-window comparison, saved-view scope, and exact request-log filters

observability

planning

/v1/apps

App attribution and metering

ecosystem

apps

/v1/apps/installs

Install, approve, and meter marketplace apps

ecosystem

apps

/v1/apps/submissions

Builder listings, runtime choices, review state

ecosystem

apps

/v1/apps/runtimes

Hosted, self-hosted, serverless, and dedicated app runtime choices

ecosystem

apps

/v1/apps/reviews

Runtime, policy, meter, and playground evidence packets

ecosystem

review

/v1/apps/readiness-preview

Simulate route, meter, policy, evidence, and runtime readiness

ecosystem

planning

/v1/apps/runtime-preview

Preview workspace file scope, sandbox continuity, network policy, routes, credits, and capacity

ecosystem

planning

/v1/apps/launch-packets

Save a reviewed app launch packet without publishing it

ecosystem

review

/v1/workspaces

Members, roles, projects, keys

account

admin

/v1/workspaces/:id/change-preview

Preview affected keys, apps, routes, budgets, and capacity before assignment

account

planning

/v1/workspaces/:id/menu

Workspace switcher, credits, labs, logs, and route shortcuts

account

console

/v1/console/surfaces

Dashboard, model hub, playground, endpoints, storage, and runtime shortcuts

account

console

/v1/billing/exports

Credit, invoice, coupon, app, and partner settlement exports

billing

finance

/v1/policies

Retention, region, vendor rules

enterprise

policy

/v1/guardrails

Prompt, output, and sensitive data checks

enterprise

safety

/v1/gpu-lanes

Reserved and dedicated capacity

compute

capacity

/v1/capacity/promotion-preview

Compare completed demand, utilization, traffic shape, runtime fit, region evidence, and normalized exposure

compute

planning

/v1/inference/endpoints

Serverless and dedicated endpoint inventory

compute

capacity

/v1/inference/endpoints/:id/health

Queue depth, warm cache, quota, and saturation

compute

routing

/v1/inference/endpoints/:id/promotion

Route evidence, health sample, policy, and budget handoff preview

compute

planning

/v1/inference/usage

Serverless and dedicated usage by model, key, region, and time window

compute

billing

/v1/playground/history

Prompt tests, multimodal runs, route decisions, and cost traces

developer

console

/v1/observability/destinations

Manage log, trace, and warehouse sinks

operations

control

/v1/observability/events

Route decisions and usage logs

operations

trace

/v1/observability/broadcasts

Fan out route traces

operations

broadcast

API tester

Run a docs request locally

This tester changes the route and generates the exact response shape developers expect.

Request

TypeScript
{
  "model": "aurona/auto",
  "messages": [{ "role": "user", "content": "Compare provider routes." }],
  "provider": {
    "only": ["approved-frontier", "approved-fallback"],
    "require_parameters": true
  },
  "route": {
    "data_policy": "zero_retention",
    "service_tier": "priority",
    "budget_key": "docs-demo",
    "return_trace": true
  },
  "observability": {
    "broadcast": ["usage_logs", "webhook", "billing_export"]
  },
  "guardrails": {
    "prompt_injection": "block",
    "sensitive_data": "metadata_only"
  }
}

Response

idle
{
  "id": "chatcmpl_demo_aurona_auto",
  "model": "aurona/auto",
  "sdk": "TypeScript",
  "policy": "zero_retention",
  "credits": "0.184",
  "latency": "640ms",
  "service_tier": "priority",
  "broadcast": "queued",
  "trace": "returned"
}

Authentication

Keys should map to projects, budgets, and routes.

A serious AI API needs more than one secret. Aurona should make environment scoping, budget attribution, and webhook verification explicit.

Key type

Where used

Purpose

Bearer token

Authorization header

server-side API calls

Management key

admin API and automation

provision projects, budgets, routes, and logs

Project key

app or environment scope

separate dev, staging, prod

Route key

alias and budget scope

share one route without exposing the full workspace

Endpoint key

serverless or dedicated endpoint scope

test and meter GPU-backed inference lanes

Policy key

policy rules and review exports

manage rule priority, actions, and evidence packets

Provider key

encrypted vault entry

route approved calls through customer-owned credentials

Budget key

usage and spend scope

customer, team, app, route

Webhook secret

signature verification

billing and route events

Provider keys

Support customer-owned credentials without losing route control.

Some customers want to bring an existing provider account. Aurona can still own the model catalog, policy checks, budget keys, traces, app attribution, and fallback rules.

Mode

Credential source

Best use

Aurona-managed

Aurona provider account

default model access, credits, route margin

Customer vault

workspace encrypted key

BYOK routing with Aurona policy, logs, and budget controls

Hybrid

Aurona plus customer key

use customer-preferred provider first, keep approved fallback

Private lane

dedicated endpoint or GPU lane

enterprise or compute-partner capacity with token settlement

Provider-key route

provider.key_mode
Show code21 lines
{
  "model": "aurona/auto",
  "provider": {
    "key_mode": "customer_vault",
    "vault_key": "pv_customer_openai",
    "order": ["preferred-frontier", "approved-fallback"],
    "only": ["preferred-frontier", "approved-fallback"],
    "require_parameters": true,
    "allow_fallbacks": true
  },
  "route": {
    "service_tier": "priority",
    "budget_key": "workspace-prod",
    "cost_quality_tradeoff": 5,
    "return_trace": true
  },
  "observability": {
    "destination": "workspace-usage-warehouse",
    "logs_shortcut": true
  }
}

Provider vault route preview

Review credential routing without handling a secret.

The proposed Aurona contract resolves masked workspace scope, credential readiness, priority, approved fallback, AI Token context, and GPU capacity into a deterministic planning receipt. It never accepts or returns raw provider credentials.

Layer

Evidence

Planning use

Scope

workspace, project key, route, model

Exclude credentials before route ordering

Credential evidence

masked identity, enabled state, expiry, rate-limit state

Show whether customer supply is usable without returning a secret

Priority

customer first, managed first, customer only

Make the attempt order explicit

Fallback

approved managed supply or fail-closed hold

Keep policy behavior visible when customer supply is unavailable

Capacity

serverless pool or reserved review

Connect the route plan to GPU-backed inference capacity

Receipt

scope, attempts, reason, AI Token context

Return deterministic planning evidence without changing traffic

Provider vault route preview

POST /v1/provider-keys/route-preview
Show code24 lines
const preview = await fetch(
  "https://api.aurona.ai/v1/provider-keys/route-preview",
  {
    method: "POST",
    headers: {
      Authorization: "Bearer <management-key>",
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      workspace_id: "workspace-aurona-launch",
      project_key_id: "ak_prod_masked",
      route: "aurona/auto",
      key_mode: "customer_first",
      credential_ref: "pv_masked_7Q2",
      fallback_allowed: true,
      capacity: "serverless"
    })
  }
);

// Proposed planning-only contract. It returns masked scope, health, priority,
// fallback, capacity, stages, and a stable receipt. It does not create, inspect, rotate, or delete a credential;
// route live traffic; debit AI Tokens;
// or allocate GPU capacity.

Fusion route config

Configure multi-model routing without rewriting apps.

Developers can keep one API call while product, finance, and security teams tune route behavior behind the scenes.

Fusion object

fusion + provider
Show code19 lines
{
  "model": "aurona/fusion",
  "fusion": {
    "panel": [
      "openai/gpt-5.5-pro",
      "anthropic/claude-sonnet-4.6",
      "google/gemini-3.5-flash"
    ],
    "judge": "aurona/judge-premium",
    "return_analysis": true
  },
  "provider": {
    "sort": "throughput",
    "allow_fallbacks": true,
    "data_collection": "deny",
    "cost_quality_tradeoff": 7,
    "session_id": "fusion-review-2026-07"
  }
}

Route options

Document the controls that make Aurona different.

The docs should make token credits, policy routing, provider selection, and GPU capacity feel concrete.

Field

Values

Purpose

model

aurona/auto, aurona/fusion, exact model id

Selects the route or model

route.objective

cost, quality, latency, balance, policy

Chooses a reusable routing objective before provider scoring

route.allowed_pool

workspace-approved, workspace-private-gpu

Limits model selection before scoring and fallback

route.return_decision_receipt

true, false

Returns workload class, objective, eligible pool, selected route, fallback, capacity, and simulated estimate

provider.sort

price, throughput, latency

Chooses provider objective

catalog.sort

price, context, throughput, latency, popularity, newest

Orders discovery results before route evaluation

catalog.output_modalities

text, image, audio, embeddings

Keeps multimodal discovery on one catalog surface

catalog.supported_parameters

tools, reasoning, structured_outputs

Returns only models that can honor required request controls

catalog.data_use_policy

no_contribution_only or provider_declared

Keeps provider-declared data use inside explicit workspace review

catalog.access_policy

standard_only or review_required

Fails closed on gated supply before an alias resolves

protocol.input

chat_completions, responses, anthropic_messages, video_job

Pins the incoming request contract before route selection

protocol.adapter_policy

preserve_only, reviewed_adapter

Holds the request when translation would lose a required feature

evaluation.suite_id

workspace evaluation id

Pins every route candidate to the same application cases and assertions

evaluation.release_gate

quality, tools, policy, credits, latency

Returns review evidence before a route alias or GPU lane is promoted

provider.order

approved provider IDs

Sets preferred provider order

provider.only

approved provider IDs

Limits a request to a specific vendor set

provider.key_mode

aurona, customer_vault, hybrid

Chooses whose provider credential funds the request

provider.require_parameters

true, false

Skips providers that cannot honor requested tools or schemas

allow_fallbacks

true, false

Controls substitution on failure

cost_quality_tradeoff

0-10

Tunes the route between cheapest healthy model and highest-quality answer

session_id

string

Keeps multi-turn agent work sticky to a compatible provider when useful

session.continuity

reuse, switch, hold

Returns whether the next turn can keep its route after task, health, policy, AI Token, and capacity checks

session.capacity_preference

serverless, reserved

Lets continuity planning include a GPU lane review without allocating capacity

data_policy

zero_retention, regional, private

Controls prompt handling

service_tier

shared, priority, serverless, reserved, dedicated

Chooses throughput and capacity class

compute_lane

serverless, hosted, gpu-fast, dedicated

Controls capacity path

endpoint.type

serverless, dedicated, private

Chooses managed pool, tenant lane, or private endpoint

endpoint.health_policy

queue_depth, warm_cache, quota, saturation

Lets the router avoid unhealthy capacity

endpoint.runtime

vLLM, TensorRT-LLM, container, partner

Describes the serving stack for GPU-backed routes

policy.rules

priority list

Applies workspace, route, app, or customer rules before provider selection

policy.action

block, mask, warn, reroute, review

Controls the outcome when a policy or budget rule matches

usage.filter

time range, timezone, granularity, key, model, route

Keeps billing, logs, and capacity reports aligned

metadata.customer_id

string

Connects usage to customer, app, account, or procurement labels

metadata.trace_level

none, summary, full

Controls how much route evidence is returned or exported

metadata.export

logs, billing, security, partner

Routes approved metadata to the right operating system

playground.history

save, redact, export

Captures approved prompt tests with route and credit traces

ranking_scope

model, route, app, gpu_ready

Exports the comparison table behind a route decision

app_id

string

Attributes usage to a marketplace or private-catalog app

app.launch_profile

coding, research, media

Selects the evaluation and route evidence expected for an app

app.catalog_scope

public, private

Adds the appropriate listing or workspace-policy review gate

app.workload_shape

bursty, sustained

Maps the preview to serverless, reserved, or dedicated capacity

budget_key

string

Maps usage to a credit budget

budget.interval

daily, weekly, monthly, lifetime

Controls which workspace ceiling is checked before execution

batch.completion_target

flexible, standard, urgent

Maps offline work to a simulated serverless, burst, or reserved capacity preference

batch.endpoint

chat/completions, responses, messages, embeddings

Preserves the application's request contract while the queue plan is evaluated

batch.credit_ceiling

AI Token amount

Holds a batch plan for review when its simulated estimate exceeds the workspace ceiling

batch.result_mode

per_request

Keeps row-level outcomes inspectable so a held item does not erase eligible results

batch.return_plan_receipt

true, false

Returns request volume, estimate, budget gate, capacity lane, and lifecycle without submitting work

reconciliation.return_preview

true, false

Returns delivery state, attempts, provisional lines, and a planning receipt without changing billing policy

reconciliation.usable_output

observed, not observed

Separates delivery evidence from provider work before ledger review

observability.destination

webhook, warehouse, trace sink

Sends approved route events to a managed destination

return_trace

true, false

Returns route decision metadata

Route objectives

Let teams choose the routing intent before model selection.

A useful router should expose simple objective presets, then translate them into provider order, service tier, policy checks, endpoint health, and credit ceilings.

Objective

Optimizes for

Best use

cost

lowest healthy credit exposure

support bots, batch jobs, broad experimentation

quality

highest route score and model fit

research, coding agents, executive answers

latency

fastest healthy provider and endpoint

interactive chat and customer-facing copilots

balance

weighted cost, quality, and latency

default production traffic and app backends

policy

retention, region, provider, and app rules first

enterprise, regulated, and private-catalog traffic

capacity

endpoint health and GPU lane availability first

serverless bursts, reserved lanes, and dedicated endpoints

Session continuity

Keep multi-turn routes stable, but never unconditional.

The proposed Aurona preview reuses a compatible session route only while task fit, endpoint health, workspace policy, AI Token budget, and GPU capacity preference remain eligible. Every result is planning-only and does not pin production traffic.

Decision layer

Evidence

Planning use

Session context

session id, prior task, prior route

carry useful continuity evidence without treating the prior route as permanent

Task check

same workload or changed workload

reuse only when the next turn still fits the route

Health check

healthy or degraded

switch before a failed or saturated route breaks the session

Policy gate

unchanged or review required

hold before execution when workspace rules need another decision

AI Token gate

open or ceiling reached

stop the next turn before new credits are consumed

Capacity evidence

serverless or reserved GPU review

preview the lane transition without reserving capacity

Session continuity preview

POST /v1/routes/session-continuity-preview
Show code25 lines
const continuity = await fetch(
  "https://api.aurona.ai/v1/routes/session-continuity-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      session_id: "agent-review-42",
      prior: { task: "code", route: "aurona/code-balanced" },
      next: { task: "analysis" },
      evidence: {
        route_health: "healthy",
        policy_state: "unchanged",
        budget_state: "open",
        capacity_preference: "serverless"
      },
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning-only contract. It returns reuse, switch, or hold evidence.
// It does not pin production traffic, debit AI Tokens, or allocate GPU capacity.

Model shortlists

Filter incompatible supply before ranking the rest.

Aurona shortlist planning combines modeled catalog evidence with policy, lifecycle, economics, performance, and GPU readiness. The result is a simulated route draft, not a production change or binding price.

Decision layer

Signals

Planning use

Objective

balanced, quality, cost

Select the decision lens before candidates are scored

Evidence

quality, intelligence, design fit, usage

Keep benchmarks and market demand visible without treating either as absolute truth

Economics

input, output, request, media, cache

Estimate the full AI Token exposure for the intended workload

Performance

latency and throughput bands

Separate interactive, agent, batch, and media routes

Hard requirements

retention, region, parameter, lifecycle, capacity

Exclude incompatible supply before weighted scoring

Route draft

alias, fallback, budget, service tier

Turn the winning shortlist row into a simulated production plan

Workload evaluations

Use repeatable application evidence before promoting a route.

The proposed Aurona evaluation contract pins candidate routes to the same workload cases, checks tool behavior and policy fit, estimates AI Token exposure, and previews the capacity path without changing production traffic.

Evaluation layer

Evidence

Promotion use

Suite

saved prompts, expected behavior, tool assertions

measure the application workload instead of a generic benchmark

Pinned run

route, model version, harness, parameters

keep comparisons attributable to the candidate being tested

Release gates

quality floor, case pass rate, tool accuracy

block regressions before a route alias changes

AI Token evidence

suite estimate, candidate total, ceiling

compare quality and operating cost on the same workload

Policy evidence

workspace, region, provider, retention

exclude routes that cannot satisfy the launch contract

Capacity handoff

serverless, reserved, dedicated

promote sustained eligible traffic into a GPU-backed lane

Evaluation run preview

/v1/evaluations/:id/runs
Show code25 lines
const evaluationRun = await fetch(
  "https://api.aurona.ai/v1/evaluations/support-agent-v3/runs",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      candidates: ["aurona/agent-balanced", "aurona/auto-fast", "aurona/private-gpu"],
      pin: { harness: "support-agent-v3", parameters: true },
      release_gate: {
        minimum_quality: 84,
        require_tool_assertions: true,
        maximum_ai_tokens: 40,
        require_workspace_policy: true
      },
      traffic_forecast: { monthly_requests: 1500000 },
      mode: "preview"
    })
  }
).then((response) => response.json());

// Preview-only contract: returns case evidence, failed gates, AI Token estimate,
// and a serverless, reserved, or dedicated capacity recommendation.

Router metadata

Return enough evidence to explain every route decision.

Aurona route metadata should help developers debug provider choice while giving finance, security, app, and capacity owners the fields they need for exports.

Metadata

Fields

Used by

provider.selected

provider id, route id, service tier

developer trace and support review

session.continuity

reuse, switch, hold, prior route, selected route

multi-turn route review and agent debugging

fallback.reason

timeout, rate limit, policy, price, capacity

operations and route tuning

cache.trace

affinity, key hash, read, write, miss reason

cost review and cache-aware route tuning

policy.action

matched rule, scope, action, owner

security and enterprise approval

credit.trace

input, output, Fusion, app, endpoint, coupon

billing and finance export

endpoint.health

queue depth, warm cache, quota, saturation

serverless and dedicated routing

endpoint.promotion

route evidence, policy review, budget window, owner

simulated serverless-to-capacity handoff

app.attribution

app id, install id, customer id, meter

marketplace and private catalog settlement

metadata.labels

environment, region, procurement, project

search, reporting, and account operations

workload.label

dimension, value, label-set id, sampled state

Activity, Logs, finance, and capacity review

Tool route evidence

Review endpoint reliability before tool traffic runs.

The proposed Aurona preview filters endpoints by tool-schema and workspace capacity requirements, then ranks simulated validity, workload evaluation, throughput, and AI Token evidence with a visible fallback.

Stage

Evidence

Planning use

Capability gate

tool schema, parallel execution, supported controls

exclude incompatible endpoints before scoring

Workspace gate

approved pool and capacity scope

keep public, reserved, and dedicated supply inside policy

Reliability evidence

schema-valid request rate and workload evaluation confidence

compare tool behavior instead of model popularity alone

Operating evidence

throughput index, AI Token estimate, capacity class

choose a balanced, reliability, or efficiency objective

Fallback evidence

second eligible endpoint and ordered review stages

make recovery reviewable before traffic runs

Planning receipt

simulated result, selected endpoint, fallback, evidence, and scope

save a decision without changing traffic, billing, or capacity

Tool route evidence preview

POST /v1/routes/tool-evidence-preview
Show code21 lines
const toolRoutePreview = await fetch(
  "https://api.aurona.ai/v1/routes/tool-evidence-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      route: "aurona/agent-balanced",
      workload: "support-actions",
      tools: { schema_contract: "strict", parallel: true },
      objective: "reliability",
      capacity_scope: "workspace-approved",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed preview contract. Returns simulated endpoint evidence and a fallback.
// It does not move traffic, debit credits, change billing, or allocate GPU capacity.

Policy rules

Rules should act before spend, data, or provider exposure happens.

Aurona can make enterprise controls operational by evaluating route, budget, region, app, prompt, and output rules before a request reaches a model.

Rule type

Match surface

Actions

Owner

Workspace rule

member, key, project, app

block, mask, warn, reroute

team governance

Route rule

model, provider, service tier

pin, fail closed, downgrade, escalate

routing control

Budget rule

customer, app, route, key

warn, block, require approval

finance control

Data rule

prompt, response, attachment

mask, redact, route privately

privacy review

Region rule

workspace, customer, endpoint

allow, deny, reroute

residency fit

App rule

install, runtime, meter

hold, approve, export

marketplace operations

Route eligibility

Preview the effective pool before a request spends credits.

The proposed Aurona preview intersects workspace, member, and API-key scope, then applies AI Token budget, retention preference, and GPU capacity readiness. More restrictive scopes win; an empty pool or reached ceiling returns a reviewable result without selecting a route.

Decision layer

Inputs

Preview result

Workspace scope

approved providers, models, regions, and capacity

establish the broadest available pool

Member scope

role, team, project, and app policy

narrow the workspace pool for one operator

API-key scope

route, provider, budget, and environment

apply the most specific request credential limits

AI Token budget

used, ceiling, interval, and hard-stop state

block selection before provider or GPU assignment

Retention preference

standard or zero-retention planning attribute

remove incompatible simulated routes

Region execution path

gateway processing, model inference, and tool execution evidence

fail closed when any stage lacks the reviewed region

Capacity preference

shared, serverless, reserved, or dedicated

return only operationally suitable lanes

Decision receipt

effective pool, selection, fallback, and ordered stages

give developers and reviewers the same evidence

Eligibility preview

/v1/routes/eligibility-preview
Show code27 lines
const eligibility = await fetch(
  "https://api.aurona.ai/v1/routes/eligibility-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      workspace: "enterprise-prod",
      member: "support-platform",
      api_key: "support-prod-route",
      route: "aurona/auto",
      retention_preference: "zero_retention_preferred",
      execution_region: "eu_review",
      require_region_evidence: ["gateway", "inference", "tools"],
      capacity_preference: "dedicated_or_serverless",
      budget: { used: 58, ceiling: 80, interval: "month" },
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed preview response: effective providers, eligible Aurona routes,
// selected route, fallback, regional execution evidence, capacity class,
// budget result, and ordered stages. This is not a residency guarantee.
// An empty pool returns policy_review; a reached ceiling returns budget_blocked.

Workspace change preview

Review the blast radius before policy assignment.

The proposed Aurona preview resolves a workspace change against current keys, private apps, route aliases, budgets, and GPU capacity paths. It returns affected object identities and reasons while keeping every production control unchanged.

Stage

Evidence

Read-only result

Proposed scope

providers, budget window, capacity preference

A read-only change set

Dependency scan

API keys, private apps, route aliases, GPU paths

Stable object identities and current bindings

Effective diff

unchanged, budget recheck, provider removed, capacity mismatch

Reasons per affected object

Decision boundary

safe to review or hold for owner review

No assignment, credential, billing, or capacity mutation

Service tiers

Make throughput and capacity explicit in the API.

A modern router should let teams choose the capacity class that matches the risk of the workload, from shared model access through dedicated GPU-backed routes.

Tier

Capacity path

Best for

shared

standard public model pool

experiments, prototypes, low-risk app traffic

priority

preferred provider order and health gates

production apps that need steadier latency

serverless

managed endpoint pool with autoscaling

variable workloads, app bursts, and endpoint trials

reserved

reserved throughput or GPU lane

high-volume routes with budget forecasts

dedicated

private model, tenant, or regional lane

regulated, private, or latency-sensitive workloads

sovereign

approved regional providers and private capacity

public sector, regulated, or residency-bound workloads

Endpoint lifecycle

Give teams a clear path from tests to dedicated capacity.

The same route can move from playground validation into serverless inference, reserved throughput, dedicated endpoints, or partner GPU lanes as volume and risk change.

Stage

Use

Aurona control

Playground

prompt tests and parameter exploration

save approved history with cost and route trace

Serverless

managed pool for variable traffic

watch queue depth, warm cache, quota, and model health

Reserved

predictable route or app volume

attach budget forecast and throughput target

Dedicated

tenant or private endpoint

pin model, region, network, and service owner

Sovereign

regulated regional route

require approved vendor, region, and audit export

Partner lane

qualified GPU supplier capacity

settle usage through capacity and credit ledgers

Usage settlement preview

Translate mixed usage without erasing its source evidence.

The proposed Aurona contract preserves each native meter, models an AI Token reservation, and reconciles only after complete provider-reported final usage arrives. Missing evidence leaves the wallet unchanged. All values are simulated and non-binding.

Evidence

Returned fields

Planning use

Native meter

input, output, cached, reasoning, request, image, megapixel, second

Preserve the provider-reported unit and billable identity

Reservation

request estimate + modeled translation weight + service tier

Hold a visible AI Token planning amount before execution

Final evidence

complete provider-reported usage lines

Fail closed when any required native meter is missing

Reconciliation

reservation compared with final measured usage

Explain the difference without silently rewriting the source meter

Capacity context

shared, priority, batch, or dedicated review

Carry economics toward GPU planning without allocating capacity

Planning receipt

profile, tier, evidence state, native line values

Make repeated read-only previews deterministic and reviewable

Settlement plan

POST /v1/credits/settlement-preview
Show code25 lines
const settlement = await fetch(
  "https://api.aurona.ai/v1/credits/settlement-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      workload: "image",
      service_tier: "priority",
      usage_state: "provider_final",
      native_meters: [
        { billable: "input_reference", unit: "image", estimate: 2, final: 2 },
        { billable: "output_image", unit: "image", estimate: 8, final: 8 },
        { billable: "output_pixels", unit: "megapixel", estimate: 16, final: 18 }
      ],
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning response: native meter lines, modeled AI Token reservation,
// final-usage evidence state, reconciliation preview, capacity context, and receipt.
// It does not debit a wallet, publish a rate card, or reserve GPU capacity.

Credit headroom preview

Resolve credit and lane pressure before admission.

The proposed Aurona contract reviews available AI Tokens, a workspace reserve, request exposure, current in-flight demand, and modeled route capacity before returning admit, queue, or review. All values are simulated and non-binding.

Evidence

Returned fields

Planning use

Credit evidence

available balance, workspace reserve, request estimate

hold before admission when spendable headroom is incomplete

Demand evidence

in-flight and requested concurrency

separate request pressure from the credit balance

Capacity evidence

shared, priority, or reserved-review lane

return a modeled ceiling without promising production throughput

Outcome

admit, queue, credit review, or evidence review

make the next operational step explicit and fail closed

Planning receipt

inputs, outcome, capacity class, simulated state

save a deterministic review record without changing traffic or billing

Admission plan

POST /v1/credits/headroom-preview
Show code23 lines
const headroom = await fetch(
  "https://api.aurona.ai/v1/credits/headroom-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      available_ai_tokens: 120,
      workspace_reserve: 20,
      in_flight: 8,
      requested_concurrency: 6,
      estimated_ai_tokens_per_request: 2,
      capacity_class: "shared",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning response: spendable credit headroom, modeled lane headroom,
// admit/queue/review outcome, ordered evidence, and deterministic receipt.
// It does not debit credits, set a production limit, or allocate capacity.

Credit ledger

Every request should create explainable billing events.

Credits become trustworthy when developers can see which tokens, media units, judge passes, route fees, and GPU lanes created the total.

Ledger field

Measured from

Settles

model.input_tokens

prompt and context tokens

provider model cost

model.output_tokens

completion tokens

provider model cost

fusion.judge_tokens

comparison and scoring tokens

Fusion orchestration

media.units

image, audio, video outputs

multimodal billing

gpu.capacity

serverless, reserved, or dedicated lane usage

compute settlement

credit.coupon

starter, promo, or enterprise credit adjustment

account balance

account.order

top-up, invoice, or committed-credit workflow

account operations

aurona.fee

routing, policy, trace, settlement

platform revenue

Credit accounts

Separate paid balances, credits, coupons, and settlement lines.

A useful AI Token account view should keep promotional credits, top-ups, commitments, app settlement, endpoint usage, and partner capacity visible as distinct ledger lines.

Account line

What it represents

Where it applies

Paid balance

customer-funded AI Token credits

model tokens, Fusion, apps, endpoints, and capacity

Starter credit

trial or onboarding allocation

developer evaluation without publishing a binding rate card

Coupon credit

promo or enterprise adjustment

separate ledger line with source, owner, and expiration state

Top-up order

self-serve or invoice request

status, amount, workspace, tax, and payment path

Committed credits

contracted drawdown plan

budget forecasts, invoices, and procurement review

Partner settlement

compute or app supplier line

route, app, endpoint, region, and approval packet

Budget intervals

Budgets should enforce before credits are consumed.

Aurona can make AI Token governance concrete by checking interval budgets, route ceilings, and app or customer caps before model, Fusion, or GPU-backed requests run.

Budget

Reset or scope

Why it matters

Daily

resets each workspace day

stop runaway spend from a single broken job or agent loop

Weekly

resets each operating week

smooth campaigns, app launches, evals, and team sprints

Monthly

resets each billing month

align credit drawdown with finance review and invoices

Lifetime

does not reset

hard cap trials, pilots, grants, and procurement-approved experiments

Route ceiling

checked per request

block expensive panels, premium models, or GPU lanes before execution

App/customer cap

checked per attribution key

keep marketplace installs and customer workspaces inside budget

Batch jobs

Plan asynchronous work against credits and capacity.

Aurona's proposed batch contract preserves Chat Completions, Responses, Messages, or Embeddings as an explicit request shape, separates offline work from interactive traffic, previews its AI Token ceiling, and returns a GPU-backed queue preference plus row-level lifecycle evidence before submission. Individual holds remain inspectable without erasing eligible rows. The contract and values shown here are simulated; planning does not submit work or allocate capacity.

Control

Example

Why it matters

Workload

text requests or embeddings

keep offline work separate from interactive and media routes

Protocol

Chat Completions, Responses, Messages, Embeddings

preserve the application contract instead of flattening every batch into one shape

Planning gate

request count, completion target, AI Token ceiling

review exposure before a queue accepts work

Capacity preference

serverless batch, burst GPU, reserved GPU

match flexible or urgent work to an explicit lane

Lifecycle

validated, queued, running, reconciling, completed

make asynchronous progress and partial failures inspectable

Row outcomes

completed or held for review

keep eligible results moving while individual rows remain independently inspectable

Plan receipt

stable simulated identifier and decision fields

save evidence without submitting work or allocating capacity

Batch plan preview

POST /v1/batches?plan_only=true
Show code22 lines
const plan = await fetch(
  "https://api.aurona.ai/v1/batches?plan_only=true",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      endpoint: "/v1/responses",
      route: "aurona/text-balanced",
      request_count: 2400,
      completion_target: "standard",
      credit_ceiling: 5,
      capacity_preference: "burst_gpu",
      result_mode: "per_request",
      return_plan_receipt: true
    })
  }
).then((response) => response.json());

// Proposed planning contract. plan_only does not submit work or allocate capacity.

Capacity promotion preview

Move from shared APIs to reserved capacity with completed evidence.

Aurona's proposed preview compares one native workload track across a completed observation window, useful utilization, traffic shape, model-runtime fit, target-region evidence, and normalized AI Token exposure. Missing evidence fails closed. The preview does not reserve GPUs, alter traffic, promise availability, publish pricing, or change billing.

Evidence

Example

Planning use

Completed window

14 or 30 observed days

Avoid promoting a short-lived launch spike as sustained demand

Native workload

routed tokens, audio minutes, or completed media seconds

Keep unlike operating units separate before comparison

Traffic shape

steady, variable, or bursty

Keep irregular demand in a flexible lane when reservation would create idle exposure

Useful utilization

observed percent plus workload-specific review floor

Compare productive capacity instead of raw GPU ownership

Runtime and region

verified model stack plus available target-region supply

Fail closed before a capacity plan reaches commercial review

Exposure comparison

normalized serverless and reserved AI Token indices

Model direction without publishing a rate, discount, or invoice

Planning receipt

normalized evidence and review outcome

Keep repeated previews deterministic without reserving GPUs

Capacity promotion plan

POST /v1/capacity/promotion-preview
Show code23 lines
const preview = await fetch(
  "https://api.aurona.ai/v1/capacity/promotion-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      workload: "agent-fleet",
      evidence_window: { completed_days: 30 },
      native_measure: "routed_tokens",
      traffic_shape: "steady",
      useful_utilization_percent: 68,
      runtime_evidence: "verified",
      region_evidence: "available",
      return_planning_receipt: true
    })
  }
).then((response) => response.json());

// Proposed planning contract. Indices are simulated and normalized.
// This does not reserve GPUs, change routes, publish rates, or debit AI Tokens.

Workspaces

Account structure should match how AI products scale.

Aurona workspaces separate personal builders, company teams, enterprises, and compute partners while keeping one control model for keys, routes, budgets, logs, credits, and invoices.

Workspace

Controls

Best use

Personal

owner, project keys, starter credits

developer trials and prototypes

Company

members, roles, invoices, budgets

team production traffic and app launches

Enterprise

SSO, audit, ZDR policy, procurement

governed model access and private routes

Compute partner

capacity profile, DD evidence, settlement

GPU supply mapped into Aurona routes

Agent connector

Coding tools should see live platform context safely.

Aurona can expose a scoped connector for model discovery, ranking checks, credit visibility, docs lookup, and safe route tests while keeping production API calls on the normal API.

Surface

Returns

Developer value

Catalog

models, context, modality, price class, route score, lifecycle

choose a model without leaving the coding tool

Rankings

usage, latency, spend, task mix, app, and GPU-ready tables

compare current routes before migration

Credit state

balance, coupons, budget interval, and route ceiling

avoid accidental spend during development

Safe test call

scoped prompt test with redaction and trace

validate a model route before production code changes

Docs lookup

quickstarts, examples, errors, and endpoint schemas

answer integration questions from the editor

Workspace scope

project, key, role, and approved surfaces

keep connector sessions bounded to the current workspace

Agent policies

Route multi-step work with explicit limits.

The Aurona Agent Route Lab turns coordinator choice, worker choice, allowed tools, child tasks, AI Token exposure, and GPU-backed capacity preference into one simulated policy contract.

Control

Example

Purpose

Coordinator route

aurona/agent-balanced

plans work and reconciles worker results

Worker route

aurona/worker-value

handles bounded child tasks on an approved route

Allowed tools

catalog, docs, safe test, app search

limits which platform surfaces the run may call

Child-task limit

1, 3, 5, or 8

caps delegation depth before execution

Run credit ceiling

workspace-defined AI Token value

stops the preview when the run reaches its budget

Capacity preference

serverless, reserved, or dedicated

maps agent work to an approved inference lane

Agent route policy

/v1/agent-policies
Show code23 lines
const policy = await fetch(
  "https://api.aurona.ai/v1/agent-policies",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      name: "aurona/agent-balanced",
      coordinator_route: "aurona/agent-balanced",
      worker_route: "aurona/worker-value",
      allowed_tools: ["catalog.read", "docs.lookup", "route.safe_test"],
      child_task_limit: 3,
      run_credit_ceiling: "0.60",
      capacity_preference: "serverless_first",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Preview response: simulated route steps, tool decisions, credit estimate,
// policy checks, and selected GPU-backed capacity class.

Workload labels

Connect what AI is doing with what it costs to run.

Aurona Workload Labels are a proposed, simulated workspace contract for attaching structured reporting dimensions to usage. Teams can preview sampling and an AI Token ceiling, then inspect the resulting labels in Logs and Activity without changing route behavior.

Control

Example

Purpose

Label dimensions

department, task, complexity, application, capacity

define the workspace reporting vocabulary

Sampling rate

10%, 25%, 50%, or 100%

preview credit exposure before broader coverage

AI Token ceiling

workspace-defined daily limit

bound the simulated labeling budget

Activity rollup

spend, tokens, requests, routes, apps, GPU lanes

compare workload shape with platform usage

Log evidence

label-set id, dimension, value, sampled state

filter and review individual simulated requests

Workload label preview

/v1/workload-labels
Show code25 lines
const labelSet = await fetch(
  "https://api.aurona.ai/v1/workload-labels",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      name: "agent-operations-view",
      workspace_scope: "aurona-launch",
      label_dimensions: {
        department: ["engineering", "product", "support", "operations"],
        task: ["agent support", "code review", "research", "media workflow"],
        capacity: ["serverless", "reserved", "dedicated"]
      },
      sampling_rate: 0.25,
      ai_token_ceiling: { amount: 12, interval: "day" },
      mode: "preview"
    })
  }
).then((response) => response.json());

// Preview-only contract: simulated labels can appear in Logs and Activity
// alongside route, app, AI Token, and GPU-capacity attribution.

Console map

The account console should match the full AI platform.

As customers move from first API call to production routes, the workspace needs clear surfaces for models, endpoint lifecycle, workflow execution, app runtime review, and finance exports.

Surface

Shows

Primary job

Dashboard

recent activity, usage trend, alerts, shortcuts

daily owner view

Model hub

catalog filters, model cards, examples, parameters

model selection and migration

Playground

prompt tests, multimodal runs, saved history

developer iteration and review evidence

Endpoints

serverless, reserved, dedicated, health, owner

capacity lifecycle management

Storage

model files, artifacts, logs, and rebuild inputs

private routes and workflow outputs

Workflow studio

node pipelines, sessions, reusable templates

GPU-backed app and media workflows

App runtime

hosted, self-hosted, serverless, dedicated

marketplace and private-catalog approval

Billing

credits, coupons, budgets, exports, settlement

finance reconciliation

Connector permissions

Live discovery and billable actions should be unmistakably different.

The proposed Aurona agent connector keeps catalog, rankings, docs, and credit checks read-only, while prompt or media tests require an explicit metered action on a separate bounded session key.

Capability

Permission

Returned or changed

Catalog and docs

read-only

models, endpoints, lifecycle, capabilities, examples, and errors

Rankings and task mix

read-only

usage, benchmarks, latency, app demand, and workload categories

Credit state

read-only

balance, budget window, coupon state, and route ceiling

Prompt test

explicit metered action

one scoped inference call with estimate, response, provider, and trace

Media test

explicit metered action

one image, audio, or video test only after cost preview

Run feedback

write to owned run

attach a category and note to a generation in the current workspace

Session key

time- and budget-bounded

separate connector access from application production keys

Runtime review

Apps should carry runtime, policy, and meter evidence into review.

Marketplace and private-catalog apps need more than install buttons. Aurona can show where an app runs, how it is metered, which routes it uses, and what evidence was reviewed.

Runtime

Control

Best for

Hosted app

Aurona-managed runtime and route meter

builder launch with simple install flow

Self-hosted app

external runtime with signed callbacks

teams that keep execution in their own account

Serverless endpoint

managed GPU-backed inference pool

bursty agents, evals, and launch experiments

Dedicated endpoint

tenant lane with pinned model and owner

enterprise or high-volume private catalog apps

Review packet

playground runs, policy actions, samples, meter

catalog approval and enterprise procurement

Version state

draft, reviewed, listed, private, deprecated

safe app lifecycle and rollback

Launch packet

route evidence, AI Token meter, policy gate, capacity sample

review before listing or reserving capacity

Proposed flow: call /v1/apps/readiness-preview while iterating, then save the reviewed result through /v1/apps/launch-packets. Saving a packet does not publish a listing or allocate GPU capacity.

App signal registration

Connect app discovery to usage and capacity evidence.

The proposed Aurona preview gives an app a stable simulated identity, then shows which public or workspace surfaces can use its route, AI Token, model-mix, and GPU capacity evidence. Workspace-only plans suppress every public surface.

Evidence

Fields

Planning use

App identity

stable app id, display name, category

join route and usage evidence without using a production API key as the app identity

Visibility

public catalog or workspace only

fail closed on public discovery, rankings, and model-app views

Usage evidence

requests, AI Tokens, model mix, time window

preview app demand without creating a real billing or ranking claim

Route context

route alias, environment, workspace

connect application demand to the model API and Fusion layer

Capacity context

serverless, reserved, dedicated review

show how app demand maps to GPU-backed inference capacity

Planning receipt

surface states, ordered checks, boundaries

save a deterministic preview without publishing or reserving capacity

App signal preview

POST /v1/apps/signal-registration-preview
Show code21 lines
const signalPlan = await fetch(
  "https://api.aurona.ai/v1/apps/signal-registration-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      app_profile: "research-copilot",
      visibility: "workspace",
      environment: "staging",
      route: "aurona/fusion-research",
      capacity: "serverless",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning contract. Returns simulated surface, usage, AI Token,
// route, and capacity evidence. It does not publish, bill, or reserve capacity.

Workspace runtime preview

Resolve execution boundaries before an agent touches a file.

Aurona can join workspace file scope, sandbox continuity, outbound network policy, model route, AI Token budget, and GPU capacity in one planning receipt. Missing scope or budget evidence fails closed.

Boundary

Evidence

Preview behavior

Workspace scope

approved file ids and app workspace

hold before route selection when the scope is missing

File boundary

isolated copies, explicit promotion review

separate durable workspace documents from temporary runtime output

Continuity

ephemeral or session reuse

declare whether later turns can see prior runtime files

Network policy

off or approved domains

keep outbound access explicit and reviewable

Route and credits

balanced or quality route, modeled AI Token budget

plan model work before any debit

GPU capacity

serverless pool or reserved review

hold restricted-budget plans without allocating capacity

Workspace runtime plan

POST /v1/apps/runtime-preview
Show code22 lines
const runtimePlan = await fetch(
  "https://api.aurona.ai/v1/apps/runtime-preview",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${AURONA_API_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      workspace_file_scope: "approved",
      continuity: "session",
      network_policy: "allowlist",
      route_objective: "balanced",
      ai_token_budget: "standard",
      capacity: "serverless",
      mode: "preview"
    })
  }
).then((response) => response.json());

// Proposed planning-only contract. It does not upload files, start a sandbox,
// debit AI Tokens, promise retention, change network policy, or allocate GPUs.

Observability

Every response should explain where it went and why.

Developers and enterprise buyers need route IDs, provider decisions, token burn, fallback events, latency, and policy outcomes.

Route trace

response.metadata.trace
Show code16 lines
{
  "route_id": "rt_aurona_fusion_01",
  "model": "aurona/fusion",
  "provider": "approved-provider",
  "service_tier": "priority",
  "key_mode": "customer_vault",
  "session_id": "agent-thread-42",
  "fallback_used": false,
  "policy": { "data_collection": "deny", "region": "us" },
  "credits": {
    "input": "0.041",
    "output": "0.088",
    "fusion": "0.055",
    "total": "0.184"
  }
}

Broadcasts

Route traces should move into the systems teams already use.

Aurona can make observability a first-class API surface by copying approved route events into logs, webhooks, warehouses, app dashboards, and finance systems.

Broadcast subscription

/v1/observability/broadcasts
Show code11 lines
{
  "sink": "usage-warehouse",
  "events": ["request.created", "provider.selected", "credit.debited"],
  "delivery": "signed_webhook",
  "filters": {
    "workspace": "prod",
    "service_tier": ["priority", "reserved", "sovereign"],
    "exclude": { "api_key": ["dev-sandbox"] }
  },
  "redaction": "metadata_only"
}

Usage logs

Route events should be useful to product, finance, and operations.

Usage logs are more valuable when they connect model choice, provider health, key mode, token burn, app attribution, and policy decisions in one exportable event stream.

Event

Fields

Audience

request.created

route, model, workspace, customer, app

developer debugging

provider.selected

provider, service tier, key mode

routing transparency

cache.affinity

scope, provider, read, write, miss reason

cost and latency review

fallback.used

previous provider, next provider, reason

operations

api_key.usage_window

key, budget, spend, requests, trend

workspace owners

endpoint.health_sample

queue depth, warm cache, quota, saturation

router and capacity owners

billing.export.created

invoice, app, credit, coupon, and partner lines

finance

route.metadata.returned

provider, fallback, policy, credit, endpoint, labels

developer support and account operations

log.filter.applied

include, exclude, request id, workspace

support and audit

credit.debited

input, output, Fusion, GPU, app fee

finance and billing

broadcast.delivered

sink, event count, retry state

SRE and analytics

guardrail.triggered

rule, category, remediation

security and product

policy.blocked

rule, region, retention, vendor

enterprise review

policy.actioned

rule, scope, action, evidence, owner

security and route approval

playground.history.saved

prompt, modality, route, cost, redaction

developer iteration

Usage reports

Billing views should reconcile model APIs and GPU endpoints.

Aurona reports should let owners filter by time window, timezone, key, model, route, endpoint mode, and export target without losing credit-ledger attribution.

Filter

Values

Why it matters

Time range

hourly, daily, monthly

compare logs, invoices, and endpoint usage consistently

Timezone

workspace local or UTC

keep finance exports and support review aligned

API key

project, route, endpoint, policy

show who created cost and which budget applied

Model and route

exact model, route alias, Fusion panel

separate raw provider cost from route value

Endpoint mode

serverless, reserved, dedicated

split variable inference from capacity commitments

Export target

billing, security, product, partner

send the right rows to each operating team

Creator and owner

member, service account, app, partner

resolve who created traffic or capacity cost

Credit source

paid, coupon, committed, partner

keep balances and adjustments auditable

Activity investigation preview

Move from a usage change to the requests behind it.

The proposed Aurona contract keeps the metric, dimensions, comparison window, saved-view scope, and exact request filters together. It joins AI Token, app, route, workspace, and GPU-lane evidence without exposing prompt content or changing live state.

Layer

Evidence

Review purpose

Metric contract

spend, requests, tokens, cache, latency, throughput

declare the measure before comparing rows

Dimension contract

workspace, app, route, model, key, workload, credit, capacity

join commercial and infrastructure evidence without flattening it

Comparison

current window plus adjacent completed window

separate a change signal from absolute usage

Request handoff

exact workspace, route, key, app, workload, credit, and lane filters

move from an aggregate anomaly to its simulated source requests

Saved-view scope

private preview or workspace preview

show intended visibility without persisting a dashboard

Planning boundary

no prompt content, no saved state, no ledger or capacity mutation

keep the website demonstration read-only

Activity investigation

POST /v1/activity/investigation-preview
Show code23 lines
const preview = await fetch(
  "https://api.aurona.ai/v1/activity/investigation-preview",
  {
    method: "POST",
    headers: {
      Authorization: "Bearer <management-key>",
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      workspace_id: "workspace-aurona-launch",
      metric: "ai_token_spend",
      dimensions: ["app", "route"],
      rollup: "day",
      compare_previous_window: true,
      selected_evidence_id: "req-aur-2048",
      saved_view_visibility: "workspace"
    })
  }
);

// Proposed read-only contract. Returns simulated aggregate rows, comparison,
// exact request-log filters, saved-view scope, ordered stages, and a receipt.
// It does not save a view, reveal prompt content, debit credits, or reserve GPUs.

Billing exports

Exports should explain credits across models, apps, and capacity.

Aurona billing exports should make it possible to reconcile usage across OpenAI-compatible model calls, Fusion runs, app meters, endpoint usage, coupons, and partner capacity.

Export

Fields

Audience

Credit ledger

paid, coupon, committed, adjustment

finance reconciliation

Model usage

fresh input, cache read, cache write, output, retry, provider

developer and cost review

Fusion usage

panel model, judge, final model, trace

route economics

App settlement

install, app, customer, builder, meter

marketplace operations

Endpoint usage

serverless, reserved, dedicated, region

capacity planning

Partner capacity

supplier, lane, health, invoice reference

compute settlement

Playground history

Prompt tests should become reviewable route evidence.

A saved playground run can carry model parameters, modality, route decision, credit estimate, redaction mode, and promotion status into docs, logs, or endpoint planning.

Stage

Use

Aurona control

Playground

prompt tests and parameter exploration

save approved history with cost and route trace

Serverless

managed pool for variable traffic

watch queue depth, warm cache, quota, and model health

Reserved

predictable route or app volume

attach budget forecast and throughput target

Dedicated

tenant or private endpoint

pin model, region, network, and service owner

Webhooks

Operational events should flow back into the business.

Aurona should notify teams when budgets, providers, apps, invoices, or policy decisions need attention.

Event

When it fires

Audience

media.job.completed

asynchronous image, video, or audio output is ready

apps, workflow studio

media.job.failed

generation stops after validation, queue, or runtime

developers, operations

credit.threshold

budget reaches warning or hard limit

finance, product

route.fallback

request switched provider or route

operations

route.broadcast

trace is copied to an approved sink

operations, finance

route.price_alert

source model price crosses route ceiling

finance, product

provider.health

latency, errors, or capacity changed

SRE

app.usage

customer or app credit event

marketplace

invoice.created

billing period closes

finance

policy.blocked

request denied by enterprise rule

security

SDKs

Meet developers where they already build.

OpenAI-compatible calls make migration easy. Native SDKs can expose Aurona-specific route, credit, app, and observability helpers.

SDK

Best for

Status

TypeScript

server apps, edge apps, agents

first-class

Python

data workflows, agents, notebooks

first-class

REST

any backend or workflow tool

stable

OpenAI SDK

swap base URL and API key

compatible

CLI

bootstrap keys, routes, and local profiles

planned

MCP

agent access to model catalog, docs, and test routes

planned

Webhooks

billing and operations events

signed

Errors

Failures should teach developers what to fix.

Clear errors make Aurona feel production-grade: credits, policy, provider capacity, and route configuration should fail with useful guidance.

HTTP

Code

Developer action

400

invalid_request

Fix malformed route, model, or parameter

401

unauthorized

Check API key and environment

402

credit_required

Add credits or increase budget

403

policy_denied

Request violates provider, region, or retention policy

429

rate_limited

Use fallback, reserved lane, or retry schedule

503

provider_unavailable

Fallback was unavailable or disabled

Guides

A documentation system for a real platform.

Aurona needs clear docs across access, credits, routing, policy, compute, apps, and observability so developers can move from first call to production with confidence.

Quickstart

Swap your base URL, add an Aurona API key, choose aurona/auto, and send your first request.

MCP and CLI

Expose live model catalog, rankings, docs, route tests, and local project setup to coding agents and terminal workflows.

Agent connector

Give coding assistants scoped access to catalog, rankings, credits, docs, and safe test routes without exposing production secrets.

Agent route policies

Preview coordinator and worker routes, tool allowlists, task limits, credit ceilings, and GPU capacity preference as one governed contract.

Coding tools

Point Codex, Claude Code, Cursor, OpenClaw, or internal agent runners at Aurona routes with scoped keys and budgets.

Model quickstarts

Browse copy-ready examples for text, image, video, audio, vision-language, embeddings, and 3D model routes.

Rankings exports

Pull usage, spend, latency, benchmark, app, and GPU-readiness tables into planning reviews and route migration notes.

Fusion routes

Run a panel of models, compare outputs with a judge, synthesize one answer, and return a route trace.

Provider routing

Sort by price, throughput, or latency while controlling fallbacks, provider order, data policy, and parameters.

Policy rules

Define priority rules that block, mask, warn, reroute, or review requests before provider selection.

Token credits

Track usage by app, customer, environment, route, model, provider, and GPU lane.

Budgets

Set prepaid credits, committed spend, alerts, hard stops, customer budgets, and route-level cost ceilings.

Budget intervals

Explain daily, weekly, monthly, lifetime, route, app, and customer budget checks before requests execute.

App marketplace

Install public or private apps with route aliases, customer budgets, app usage logs, and builder settlement events.

App launch readiness

Preview model evidence, route policy, an AI Token meter, and serverless or reserved capacity before a listing is published.

Cost simulation

Estimate model, Fusion, app, fallback, and GPU lane exposure before approving production traffic.

GPU routes

Attach reserved inference lanes and advanced GPU supply for workloads that need predictable throughput.

Inference endpoints

Register serverless pools, dedicated endpoints, queue policy, warm-cache state, quota, and runtime metadata.

Console map

Show the dashboard, model hub, playground, endpoints, storage, workflow, runtime, and billing surfaces a workspace owner expects.

Playground history

Save prompt tests, multimodal runs, model parameters, route decisions, and cost traces for later review.

Billing exports

Separate model tokens, Fusion, coupons, app settlement, endpoint usage, and partner capacity into finance-ready rows.

Guardrails

Add prompt, output, sensitive-data, vendor, and region checks as routing inputs before a provider sees traffic.

Enterprise policy

Configure retention, region, approved vendors, audit logs, SSO, spend controls, and fail-closed fallback.

Observability

Inspect route decisions, token burn, latency, provider health, fallback behavior, and customer usage.

Webhooks

Subscribe to credit thresholds, route failover, provider health, budget events, invoice activity, and policy blocks.

Advanced features

Make the docs feel like a complete developer platform.

The docs go beyond the first API call into routing, fallbacks, tools, ZDR, attribution, service tiers, AI Token controls, and GPU capacity.

Feature

API surface

Why developers need it

Protocol route preview

/v1/routes/protocol-preview

Prefer native request formats, expose reviewed adapters, and verify an ordered fallback chain under one shared request contract

Request reconciliation preview

/v1/usage/reconciliation-preview

Review usable output, provider attempts, approved fallback, auxiliary work, and provisional AI Token lines without issuing credits or changing billing policy

Media job planning

/v1/media/jobs

Validate output controls, preview AI Token exposure, choose a GPU lane, and return synchronous output or an asynchronous job receipt

Media capability discovery

/v1/media/models

Discover output types and endpoint-specific resolution, aspect ratio, reference input, format, streaming, and delivery support

Model fallbacks

fallback: auto or approved-only

Switch providers or routes when capacity, price, or policy changes

Cache-aware routing

/v1/routes/:id/cache-policy

Set cache affinity and scope, then inspect provider-reported cache evidence without bypassing route policy

Provider key vault

provider.key_mode: customer_vault

Let customers attach approved provider credentials while Aurona keeps route, policy, and usage records

Management API

/v1/management/keys

Automate project, route, budget, and usage-log administration

API key detail

/v1/api-keys/:id/usage

Show per-key charts, budget progress, log shortcut, rotation state, and workspace ownership

Router metadata

/v1/routes/:id/metadata

Return provider, fallback, policy, credit, endpoint, app, and customer evidence when requested

Policy rule engine

/v1/policy/rules

Define priority-based block, mask, warn, reroute, and review actions for production traffic

Policy action log

/v1/policy/actions

Export matched rules with scope, owner, request id, evidence, and downstream billing impact

Observability destinations

/v1/observability/destinations

Manage Datadog, Langfuse-style, warehouse, webhook, and billing-export sinks from one surface

MCP model bridge

/v1/mcp

Expose live catalog, rankings, docs, and safe test inference to coding agents

Agent connector

/v1/agent-connectors

Create scoped sessions for coding tools to read catalog, rankings, docs, and credit state

Agent route policies

/v1/agent-policies

Bound coordinator and worker routes, tools, child tasks, AI Token exposure, and capacity preference

Workload labels

/v1/workload-labels

Preview workspace-scoped dimensions, sampling, AI Token ceilings, log tags, and Activity rollups

Model capability discovery

/v1/models/capabilities

Filter by modality, parameters, region, age, provider-declared training metadata, and endpoint class

Catalog query preview

/v1/models/query-preview

Filter hard capabilities, place missing evidence last, resolve a canonical route, and return a planning receipt

Ranking feed preview

/v1/datasets/rankings-preview

Return versioned ranking rows with explicit measure, window, opt-in scope, portable citation, AI Token, GPU capacity, and trend-comparison evidence

Model shortlist planning

/v1/model-shortlists

Rank eligible routes by balanced, quality, or cost objectives after hard policy and capacity exclusions

Workload evaluation

/v1/evaluations/:id/runs

Compare pinned route candidates on application prompts, tool assertions, quality, AI Token use, policy, and capacity evidence

Model lifecycle

/v1/models/lifecycle

Expose preview state, version pinning, retirement date, replacement route, and workspace owner

CLI bootstrap

aurona login + aurona route init

Help developers configure keys, route aliases, budgets, and local profiles quickly

Coding tool setup

OpenAI-compatible endpoint config

Provide copy-ready setup for Codex, Claude Code, Cursor, OpenClaw, and internal agent runners

Model quickstarts

/v1/model-quickstarts

Group text, image, video, audio, and 3D examples by model capability

Tool calling

tools + require_parameters

Route only to models and providers that support function calling

Structured outputs

response_format + schema

Keep JSON, scoring, and app automations reliable

ZDR

data_policy: zero_retention

Restrict prompts to providers and lanes with compatible retention rules

App attribution

app_id + customer_id

Settle usage to app builders, enterprise catalogs, and customer budgets

Service tiers

service_tier + compute_lane

Choose public, priority, reserved, or dedicated capacity

Route objectives

/v1/route-objectives

Reuse cost, quality, latency, balance, policy, and capacity presets across apps

Budget intervals

/v1/workspaces/:id/budgets

Enforce daily, weekly, monthly, lifetime, route, app, and customer ceilings before execution

Cost simulator

/v1/cost-simulations

Preview model, Fusion, app, and GPU credit exposure before routing production traffic

Guardrails

/v1/guardrails

Block prompt injection, sensitive data, unsafe outputs, or disallowed vendors before provider selection

Observability broadcast

/v1/observability/broadcasts

Send route traces into logs, webhooks, finance systems, and monitoring tools

Usage logs

/v1/observability/events

Inspect token burn, app attribution, fallback, policy, provider-health events, and request-id filters

Rankings export

/v1/rankings

Fetch usage, spend, latency, benchmark, app, and GPU-readiness tables for internal review

App marketplace install

/v1/apps/installs

Attach route aliases, app meters, customer budgets, and usage-log shortcuts during install

App runtime review

/v1/apps/reviews

Attach saved playground runs, runtime metadata, policy actions, and meter evidence before listing

Workspace menu state

/v1/workspaces/:id/menu

Show credits, labs, keys, logs, routes, and billing shortcuts in the account surface

Console surface map

/v1/console/surfaces

Expose dashboard, model hub, playground, endpoints, storage, workflow, runtime, and billing shortcuts

GPU endpoint tester

/v1/gpu-lanes/:id/test

Preview queue depth, model cache, latency, and capacity class before routing production traffic

Inference endpoints

/v1/inference/endpoints

Register serverless pools, dedicated endpoints, model cache, quota, health, and runtime metadata

Endpoint lifecycle

/v1/inference/endpoints/:id/lifecycle

Move a route from playground testing to serverless pool, reserved lane, or dedicated endpoint

Billing exports

/v1/billing/exports

Separate model tokens, Fusion, coupons, app settlement, endpoint usage, and partner capacity lines

Agent runtime path

/v1/apps/runtimes

Choose hosted, self-hosted, serverless GPU, or dedicated endpoint before an app listing is reviewed

Filter grammar

logs.filter

Support include and exclude chips for model, provider, key, workspace, route, and event type

Router metadata

metadata object

Attach customer, environment, region, and procurement labels to every request

Aurona.ai

Build on the token network for AI applications, models, and compute.

API

OpenAI-compatible

Billing

AI Token credits

Capacity

GPU-backed routes