Use case
Self-Hosted AI Agents: A Platform to Run or an App to Own
Self-hosting an agent platform and owning an agent application are two different jobs. The first gives your team a builder to operate. The second gives your customers a product, and the agent is one feature inside it.
Most round-ups of self-hosted ai agents list the same handful of names, n8n, Dify, Flowise, LangGraph, and end with a Docker Compose file. That answers one question well: which builder can I run on my own server so internal automations stop passing through somebody else's cloud. If that is your question, pick one of those and close this tab. We sell a codebase, not a hosted product and not a workflow builder, and nothing below competes with them for that job.
The lists are thinner on the other reader: the developer or technical founder who wants to ship an agent as a product. That product has sign-up, workspaces, a subscription, chat history, uploaded files and a usage cap per customer. A builder you operate does not turn into that by adding a login page. This page covers the difference, what "self-hosted" still leaves outside your walls, and what you sign up to run.
Two things the phrase covers
Who talks to the agent when it is finished?
Run a self-hosted platform.
- n8n for automations with agent steps, Dify and Flowise for assembling agents and retrieval flows on a canvas
- You operate their containers and database, and build inside their editor rather than in a repository
Own an agent application.
- Accounts, organizations, billing and per-customer limits are most of the code, and the agent loop is a small part
- A framework such as LangGraph or the Vercel AI SDK gives you the loop, in code, and nothing around it
- You write and deploy an ordinary web app, which means you also own its queue, its migrations and its upgrades
The first branch is infrastructure you adopt. The second is software you write.
What self-hosted covers when the model is an API
"Self-hosted" describes where the application runs. It says nothing about where the tokens are computed. Most self-hosted AI setups, platform or application, still send every prompt to a model vendor.
An agent stack, by who actually holds it
Routes, tools, permission checks, the run loop. On your servers. Every self-hosted option gives you this.
Conversations, uploaded documents, embeddings, usage records. In your Postgres and your bucket, so retention is your decision.
Prompts, retrieved passages and tool results leave your network on every call unless you also run the model. The vendor's retention terms apply, and so does its per-token bill.
Ollama or vLLM on your own GPUs closes the last gap. You trade a per-token invoice for hardware, and smaller open models call tools less reliably, so test your own tools before committing.
If the requirement is data residency, the third row is the one an auditor asks about. Docker does not answer it.
Cost control and no platform lock-in are met by the first two rows. A rule that customer text may not reach a third party needs the fourth, and that changes your hardware budget more than your software choice.
What you take on
A hosted agent service hides a short list of chores. Self-hosting hands them to you, on either branch.
The operational bill for an agent you host
- Runs that outlive a requestEvery deploy restarts a process in the middle of a stream. Either the run is recorded somewhere durable and something finishes or fails it, or the user watches a spinner.
- A queue for slow workParsing and embedding a large PDF does not fit in an HTTP request. You need retries, backoff and a concurrency limit, or one upload starves the API.
- Tool calls that are safe to repeatA retried run calls its tools again. Anything with a side effect needs a record of what already happened.
- Secrets and spendProvider keys stay on the server, and a per-customer cap has to be enforced before the call, because the vendor bills you for a runaway loop either way.
- The database and its upgradesBackups, migrations, and a plan for changing the embedding model, since vectors from two models cannot be compared.
Starting an agent application from Aether
This is the part we make. Hype Stack is an MIT-licensed full-stack template (React, Hono, Prisma and Kysely on
Postgres), and Aether is the template we build on it for AI products. It installs a Better Auth or WorkOS starter with
organizations, a billing pack, an admin app, and pack-ai-chat: a chat workspace at /chat and a document library at
/knowledge, on the web and in the Expo app. The source is copied into your repository.
How the pack answers the checklist above
Runs and tool calls are rows
Every run and tool call is written to Postgres. A cron job named ai-chat-recovery runs every 30 seconds and finishes or fails runs whose process died. Tools go through a journal keyed by run and call id, and authorization runs again on a replay, so a cached result cannot outlive revoked access.
- ai-chat-recovery
- libs/ai-tools/define-tool.ts
Indexing runs on a queue inside Postgres
Each upload becomes one knowledge-indexing task on the template's pg-boss queue, written in the same transaction as the upload. It retries five times with backoff and resumes from the last embedded batch. The queue lives in Postgres and works inside the API process, so there is no broker or worker service to deploy.
- pg-boss
- KNOWLEDGE_INDEXING_CONCURRENCY
Spend is reserved before the call
Tokens, operations, voice minutes and storage are reserved per organization before a run and settled after it, against caps from the billing plan.
- ai-usage.prisma
- /ai-usage
The model is the foreign part, with its own account and its own bill. The pack ships provider definitions for OpenAI,
Anthropic and Google. AI_TEXT_MODEL takes a provider:model value, so the text model moves between those three with
an environment variable. An OpenAI key is the one you need to start, because voice and the default embedding model point
there. In the terms of the stack above, Aether gives you rows one and two. Row three stays with the vendor.
Retrieval does not add a vendor. Chunks are stored in Postgres with the embedding in a Float[] column on
knowledge_chunk, and ranking is one SQL query that mixes a dot product with Postgres full-text search. There is no
vector database to run and no extension to install. The cost is that the query scans a user's selected documents with no
approximate index. That suits a per-user document library. A corpus of millions of chunks would want pgvector or a
dedicated store, which you would add and operate.
Deployment is containers. Each app has its own apps/<app>/Dockerfile, and npx @hype-stack/cli deploy web puts the
frontend, admin, backend, Postgres, Valkey and a bucket on Fly.io or Railway, both of which bill you directly. The
images build anywhere Docker runs, but on your own hardware the wiring is yours: the CLI does not target a bare server.
A run is capped at five model steps in
model-response.tsand atAI_MAX_RUN_SECONDS, which defaults to 180 and cannot be set above 600. That is a chat turn with tools, not an agent that works unattended for an hour.
What is not in it
No visual workflow builder. Nobody draws an agent on a canvas here. If your users need to, a self-hosted platform is the right base.
No local model runtime. There is no Ollama or OpenAI-compatible provider in the pack. Adding one means writing a
provider definition next to the three in features/ai-chat/modules/models/providers, and if you also move embeddings to
a local model, every document has to be reindexed: search refuses to compare vectors across models and returns a 409
asking for exactly that.
No multi-agent orchestration, no MCP client, and no long-running autonomous runs. The built-in tools are web search, image generation and document search. For work that spans hours or needs human approval steps, the usual answer is a durable execution engine such as Temporal or Inngest, which are other companies' products that you run or pay for alongside this.
What you write yourself
- Your tools. A contract in the shared enums package, then an implementation with
defineChatToolthat suppliesauthorizeandexecute. The journal and the permission check on replay come with the wrapper. - The system prompt and the product around it. The pack ships a general assistant.
- Anything beyond a chat turn. Scheduled agents, background runs and approvals are yours to build, on the queue and scheduler that are already there or on an engine you bring.
If the first branch is yours, a platform will have you running this week. If it is the second, the sections below show what the Aether install puts in your repository.
What is Hype Stack?
Every product starts with the same month of work nobody pays you for: sign-up and login, teams and permissions, taking payments, notifications, an admin panel to run the business. Hype Stack is that month, already built and tested. Start from the free open-source app, add the pieces you need with one command, and keep going on the part that is actually your idea.
Everything lands as real code in your own repository, so there is nothing to rent and nothing anyone can switch off. For the engineers: React 19, Hono, Postgres, and a desktop build, typed end to end.
See it running
Templates are curated project starters built on this stack: a layout, feature packs, a custom theme, and bonus pages. The previews below are recordings of the real apps.
Two ways to install features
Same features underneath, different starting point. Either command resolves what the packs depend on, copies the source into your repository, and merges the Prisma schema.
Take a template
A landing page, a design system, a themed layout, and the features already wired into it. Rebrand it, put your product in the middle, ship.
npx @hype-stack/cli template$ hype-stack compose
Compose your own design
Your design and your choices, without rebuilding auth, billing, or notifications. Tick the packs you want and the CLI wires them into the open-source starter.
npx @hype-stack/cli composeInside the stack
What each pack gives you
Source code, not a dependency. Every pack lands in your repository across the surfaces the feature touches.
Authentication, organizations, roles, sessions, and a full admin app, powered by Better Auth on your own Postgres.
- Email & Google login
- Organizations & members
- Roles & permissions
- Admin app & dashboard
Questions, answered
More stacks
Turn your ideas into
Real applications.
Start free and own every line you ship. When you want more, one All-Access license unlocks every premium pack and template for a year.