AlevatedOS · The operating system for AI agent fleets

Your agents are running.
Do you know what they're doing?

AlevatedOS is the mission-control layer for any team running multiple AI agents in production — engineering, marketing, operations, growth. Wherever a fleet is live, this is how you operate it. One briefing. One board. One approval gate. Built for operators who run real fleets — not demos.

  • 60-second voice briefing from Iris, your AI ops lead
  • Seven-lane Bridge board across every live project
  • Keyboard-first approvals for every consequential agent action
The reality

Most teams run their agents blind.

You set up the agent. You handed it the tools. Then traffic moved on and you stopped looking. That's where the silent failures live.

01

No morning brief

You spend 20 minutes every morning piecing together where things stand from Slack, docs, and three open dashboards. By the 9am call you're already behind.

02

No approval gate

Agents take consequential actions — provisioning teammates, sharing secrets, signing off on copy — with zero human checkpoint. When it goes wrong, it goes wrong fast.

03

No early warning

A project goes quiet. You find out three weeks later, on a customer call, that the brief never landed. There's no objective signal that a fleet member has drifted off-script.

A typical operator morning

From scattered to in command in under sixty seconds.

Four moves. The same four every morning. AlevatedOS sequences them so you walk into the 9am with a current picture of every fleet member.

STEP 01

Iris briefs you like a field commander.

Sixty seconds of audio. Or read it inline if it's not that kind of morning. Iris pulls the state of every active project, scores it against the trajectory model, and gives you the four lines that matter: blocked, waiting, silent, moving. Then she collapses into a status strip until something changes.

  • OpenAI-powered TTS, walkie-talkie cadence
  • Replays on demand · keyboard SPACE to play
  • Regenerates on demand · cached between runs
TRANSMISSION 06:42 PST
I IrisAI Ops Lead 00:54
2 blocked 1 waiting 1 silent 4 moving
STEP 02

The Bridge Board · every project, every status, at a glance.

A spatial phase pipeline — configurable per team. The default lanes (Intake → Spec → Build → Review → Deploy → Verify → Live) fit engineering, marketing, and operations work alike. Every project is a token. Bright means moving. Azure glow means needs your decision. Red and sinking means blocked. Desaturated means it went quiet.

  • Rebuilt from the record on every load
  • Hover a token for the last three agent actions
  • G + B from anywhere to jump here
BRIDGE · 8 ACTIVE
Spec
Build
Review
Deploy
Verify
Live
STEP 03

Approvals Inbox · agents don't act alone on the big stuff.

New agent provisioning. Secret sharing. Phase sign-off. Anything with a real-world consequence flows through the Approvals Inbox. Keyboard-first by default. jk to navigate. e to approve. r to reject. The agent waits for you, not the other way around.

  • Inbox refreshes every 30 seconds · the agent picks your verdict up on its next 60-second heartbeat
  • Full audit trail · who approved what, when, why
  • Policy knobs: auto · draft · off · operator-only per capability
Approvals · 3 pending jk nav · e approve · r reject
SECRET ROTATION Rotate CF_PAGES_DEPLOY_TOKEN for naustar.com Quinn · 47s ago · naustar-site
AGENT PROVISIONING Provision Mira · social-content sub-agent Iris · 6m ago · fleet-wide
PHASE SIGN-OFF Phase 4 → 5 release cut for PRJ-METIS Alf · 14m ago · BLD-014
STEP 04

Trajectory Scorecard · the test that runs whether you check it or not.

Every project gets a weekly alignment score. RED AMBER GREEN. Automated. Objective. No manual entry, no operator wishful thinking. The score runs in the background and lands in the briefing when something turns.

  • Composite of phase velocity, agent activity, customer signals
  • Verdict + top blocker line + one suggested next move
  • Lands in the briefing when a verdict turns
project-cedar PRJ-CEDAR
AMBER 62/100
  • Phase 3 review approval has been open 4 days
  • Agent activity nominal · last action 11m ago
  • Suggest: nudge stakeholder in dashboard chat
trajectory · weekly run
What's inside

The operator's surface.

Six tools. One keyboard. Zero context-switching.

Voice briefing

OpenAI-powered TTS. Iris briefs you like a field commander. Plays in 60 seconds, collapses to a status bar after you’ve heard it. Regenerates when the picture materially changes.

Bridge board

Seven-lane phase pipeline · configurable per team. Rebuilt from the record on every load. Hover a token for the last three agent actions.

Approvals inbox

Human-in-the-loop for every consequential action. Keyboard-first. Per-capability policy: off · draft · auto. Full audit trail.

Trajectory Scorecard

RED / AMBER / GREEN every week. Composite of phase velocity, agent activity, customer signals. Runs whether you check it or not.

Agent fleet

Roster of every agent — model, capabilities, current session, last action. Provision new agents through the wizard with the operator gate already wired.

⌘K

Command palette

Fleet-wide command from anywhere. Route a spoken command to approvals, fleet, or a specific project.

Inside the OS

Every capability you'd expect from a real OS.

Not a wishlist. Every line below is shipping right now against live project rows in a production fleet.

RUNTIME Hermes & OpenClaw agents · side-by-side.

AlevatedOS is runtime-agnostic. Operate Hermes agents (Claude-native, SDK-built) and OpenClaw agents (your custom-framework fleet) from the same surface. Same Bridge board. Same approval gates. Same trajectory scoring. No other operator layer does this.

Hermes · Claude-native OpenClaw · custom fleet Bring-your-own · roadmap
VOICE OpenAI TTS briefings

60-second audio brief, regenerated when the picture changes, playable from anywhere.

BRIDGE 7-lane phase board

Spatial pipeline · configurable per team. Default: Intake → Spec → Build → Review → Deploy → Verify → Live.

APPROVE Keyboard-first inbox

jk navigate · e approve · r reject. Agents poll on a 60-second heartbeat.

SCORE Trajectory RAG verdict

Weekly run · phase velocity + agent activity + customer signal.

FLEET Agent provisioning wizard

Spin up a new agent with model, capabilities, and policy gates in one flow.

POLICY Per-capability autonomy

Every action is off · draft · auto. Operator-tunable, never hard-coded.

SECRETS References, never values

The registry records where a secret lives — never the secret. No route accepts one.

SILENCE Quiet-agent detection

Auto-flags a fleet member that hasn't transmitted in N hours. Surfaces in the briefing.

AUDIT Full action ledger

Who approved what, when, why. Replayable. Exportable for compliance.

⌘K Fleet-wide command

One palette routes spoken commands to approvals, fleet, or a specific project.

SPEND Cost per agent, per day

Token spend metered on the same roster as status — so a runaway agent shows up before the invoice does.

QUEUE Work queue with receipts

Dispatch → checkout → check-in, each item carrying its own event timeline and artifacts.

The record

Ask for the log. We can produce it.

Three moments from a single day of production, each one a row you could read back. Timestamps are UTC on 2026-08-24; anything in quotes is verbatim, not paraphrased.

  1. 07:42:58 07:43:05 07:43:14 16s

    Dashboard to a machine on another continent, and back.

    An item was dispatched, claimed by an agent on its own hardware, executed, and checked back in — sixteen seconds end to end, no human hand in flight. A person pressed dispatch. Nothing after that was a person.

    The honest half: this was the second attempt. The first failed 44 minutes earlier and said why, in the agent’s own words, in the same timeline. We kept both rows.

    dispatch 07:42:58 checkout 07:43:05 check-in 07:43:14 check_events seq 97–99
  2. 12 blocked

    Failure that looks like failure, on purpose.

    Twelve items stopped that day. Not one of them stopped quietly — every blocked event carries the reason that produced it. Three runs that exited zero were caught, reopened and annotated, because the agent’s own check-in admitted no code had been written.

    A representative reason, unedited: “claude -p runs from $HOME so the repo allowlist is never read — the 7m34s completion was honest-empty #2.” That is what a real one looks like.

  3. 08:10:30 08:10:44 froze

    The agent stopped and asked.

    An agent noticed work executing in a way that contradicted its last human instruction. It did not correct the system, and it did not carry on. It froze mid-run, escalated with three named options, and wrote:

    “Not modifying the daemon or the process until Andrew replies.”

    The audit log answered it: every dispatch it had flagged was operator-initiated, logged with a human as principal, under a tier the operator had explicitly chosen. The alarm was wrong. Raising it was right, and the record is what settled it.

    It also asked what those runs had shipped. The answer in the log was nothing — zero bytes persisted. We would rather publish that sentence than not have been able to answer the question.

Routing

The record shows the model that actually ran.

Routing happens at the work layer, not the request layer: a directive rides on the work item, and the agent reports back what it really invoked. Which means the interesting question isn’t what we routed — it’s what we can prove afterwards.

Every completion lands in one of three states
  • honored a directive was set, and that exact model ran
  • mismatch a directive was set and something else ran — recorded, never smoothed over
  • unreported no directive; the box default ran and the record does not name it

The coverage ratio is printed next to the table, always. Not the flattering half of it — the denominator too, at zero as readily as at full. A table built from part of the evidence must never read as the whole fleet, so the panel tells you what fraction it is before you read a single row.

The honored one, verbatim
{ "type":            "model_routing",
  "model_directive": "anthropic/claude-haiku-4-5-20251001",
  "model_arg":       "claude-haiku-4-5-20251001",
  "model_source":    "item_directive" }

Reported by the agent on check-in. Asked for, invoked, and where the instruction came from — three separate fields, so they can disagree in the open.

  • Ground truth The roster reads each agent’s model from its own session log — what it is really running, not the override an admin typed.
  • Never a guess If the reported id doesn’t parse as vendor/model, the dashboard shows “no model set” rather than inventing a plausible one.
  • Pending ≠ applied A model change that hasn’t actually taken on the box reports itself unapplied and keeps showing pending until it truly lands.
Architecture

Your machines. Your keys. Your Postgres.

The agents don’t run in our cloud. They run on hardware you control, with your own model API keys, reaching the control plane through a scoped credential you can revoke. Nothing about your fleet has to live in somebody else’s account.

The whole system, drawn honestly
Control plane
Dashboard Next.js on Cloudflare Containers, behind Cloudflare Access
Postgres Supabase, via Prisma
agk_ scoped bearer key heartbeat · skill sync · usage · config pull & apply · work loop
Yours — below this line we hold nothing
Agent box #1 your machine · your model API key
Agent box #2 your machine · your model API key
… #N one dependency-free Node file, launchd or systemd
  • The credential agk_ + 256 random bits, stored as a SHA-256 hash — the plaintext is shown once and never persisted. Scoped to named actions, TTL-able, and revocable one way: revoked stays revoked.
  • Revoke means now The credential is read from the database on every single request, so a revocation or a scope change takes effect on the next call. Granting an agent a new scope needs zero key rotation — and the grant itself lands in the audit trail.
  • Secrets stay yours The registry records where a secret lives — a 1Password item, an env var name. Never the value. There is no column and no route that accepts one.
  • Receipts Every privileged action lands in an append-only, admin-only audit trail, filterable and exportable.

Said precisely, because the distinction matters: this is instance-per-customer — your own deployment, your own database, not a tenant row in ours. One organization exists today: ours, running the eight projects above. The multi-tenant path is built but has not been exercised across two organizations, and you should know that before you plan around it.

Where this sits

Three shapes. Only one is yours.

We compare architectures, not products — on purpose. A grid of checkmarks about other companies’ software is a wall of claims we can’t verify and can’t keep current, which is exactly the kind of thing this page exists to not do. What follows is definitional: three structurally different ways to put a fleet under control. Place the vendors you’re evaluating yourself.

Structural comparison of three agent-control architectures
Request-level router sits between an app and model vendors Hosted agent platform your fleet runs in the vendor’s cloud Control plane you own this one
Where the agents run Nowhere — it forwards calls, it doesn’t run agents The vendor’s infrastructure. That is what hosted means Your machines, under your process manager
Whose model API keys Typically the router’s, with usage billed through — that is the value on offer Varies by vendor Yours. They never leave the box
The unit being governed A request A task or a run A work item — with an agent, a project, a tier and an audit row
What the routing decision can see The request: model, tokens, maybe a tag Whatever the vendor’s model of a task exposes Who is doing the work, on what, at which tier
Where the record lives The router’s logs The vendor’s database Your Postgres. You can query it without asking anyone

And the trade, since a comparison that only flatters one column isn’t worth reading: owning the control plane means operating it. You run a database and a daemon per box, and you carry your own model spend directly instead of through someone else’s invoice. If you don’t want to run infrastructure, a hosted platform is genuinely the better answer — and the first two columns are real products solving real problems, not straw men. This shape is for operators who need the record to be theirs.

Cost

Find out what an agent costs before the invoice does.

Spend is metered per agent, per day, on the same roster as status — so “which one of these is expensive?” has an answer that takes a glance instead of a billing cycle.

week of Aug 10 $1,439.53
week of Aug 17 $10.88

One agent, one model change. It was findable at a glance because the spend was attributed to the agent that caused it — rather than pooled into a single line on a bill at the end of the month.

And where a cost hasn’t been reported, the panel says “not reported” — never a confident $0.00. A fake zero is the one number that would make the whole table worthless.

Read honestly: one agent over two weeks, not a fleet average. A different agent’s $2,826.33 week the same month was a known, deliberate audit — explained in one question instead of one billing cycle. Metering began 2026-06-29, so there is nothing earlier to chart.

READY?

Run your fleet from one screen.

Early access · hand-run onboarding · 30-minute working session, not a sales call.

Request access No credit card · Reply within 24 hours
Built for operators who run real fleets

Not a deck. Live infrastructure.

8 active projects on the Bridge
15 agents on the roster · each with its own heartbeat, spend and tier
60s heartbeat — a free agent picks up work on its next poll
4 governance outcomes per tier · autonomous → blocked

Project and roster counts read from the production database on 2026-08-24, 22:33 UTC — a live count, not a lifetime total. The poll interval and the tier lattice are fixed properties of the software, not measurements. Live numbers move; we re-query them before we publish them, and we date-stamp what we quote.

  • DEPLOYMENT Running in production on The Marketing Pros’ own agency fleet: 8 active projects across a 15-agent roster, every agent carrying its own heartbeat, spend and tier. The work loop — dispatch, execute on the agent’s own machine, check back in — closed on 2026-08-24. Those eight aren’t demo rows: they are the product lines and client-delivery operations this company actually runs, which is why the dashboard has to be right.
  • STACK Next.js on Cloudflare Containers behind Cloudflare Access, Supabase Postgres, and a small dependency-free daemon on each agent’s own machine. No demo data, no canned screenshots. Every UI surface ships against real database rows.
  • GATING You choose what’s gated — per tier, per org, per project. The lattice runs autonomous → notify → require approval → blocked, an override can only ever raise oversight, and the record shows which rule fired for every action. A tier with no rule configured falls back to require approval — it never fails open. The agents that executed work on our own fleet today run at tier 1, where the rule is notify rather than sign-off. That was a recorded operator decision, and every dispatch it produced carries the line “tier-1 notify — proceeded autonomously” in the log.
  • NO SILENT EXPIRY An approval that nobody answers stops the work and says so. It does not quietly expire into a dispatch: repeated expiry trips a breaker that parks the item as blocked and escalates, rather than churning forever.
Early access · invite only

Built for operators. Not demos.

AlevatedOS is rolling out to a small group of teams already running live AI fleets — engineering, marketing, operations, anyone with more than one agent in production. Drop your work email and we'll set up a 30-minute working session — walk your morning, see if the fit is real.

  • Working session, not a sales call
  • Bring one project that's currently silent
  • If it's not a fit, we tell you on the call
  • Iris herself writes the follow-up

▸ Briefing transmitted. We'll be in touch within 24 hours.

No marketing list. No drip sequence. Iris will write you directly.