Back to guides
Setup·October 1, 2026·7 min read

AutoGPT custom API: running self-hosted AutoPilot on your own OpenAI-compatible endpoint

Four environment variables move AutoPilot chat onto your own endpoint on a self-hosted install. The block layer and the hosted service do not follow — and embeddings and memory calls land on the same URL.

Last reviewed on .

On a self-hosted AutoGPT Platform, the AutoPilot chat agent runs on any OpenAI-compatible endpoint you give it — four environment variables, set on the container. We ran v0.8.2's official single-container image that way against a recording stand-in on October 1, 2026: a chat turn and three tool round trips completed, and the workspace file AutoPilot wrote was read back through the API. The hosted agpt.co service does not take a custom endpoint, and neither does the AI Text Generator block inside agent graphs.

For what AutoGPT costs overall — hosted plans, credits, self-hosting — see AutoGPT pricing. If you are leaving AutoGPT, AutoGPT alternatives. This page is the custom-endpoint route itself.

What can use your endpoint, and what cannot

Part of AutoGPTTakes a custom endpoint?Basis
AutoPilot chat, self-hostedYes — CHAT_BASE_URLRun on v0.8.2: chat and tool calls completed
Conversation titlesYes, same endpoint and modelSeen in the run: a short non-streaming call per session
Builder dry-run simulator, onboarding extractionYes, per AutoGPT's guideDocumented; not exercised here
Graphiti memorySame base URL, its own models (GRAPHITI_*_MODEL)Seen in the run: /v1/responses plus an embedding call per turn
Marketplace semantic searchSame base URL, STORE_EMBEDDING_MODEL (must return 1,536-dimension vectors)Seen in the run: 200 embedding calls at boot
AI Text Generator block in agent graphsNo — fixed providers; only an Ollama-native host is settablev0.8.2 source
Hosted agpt.coNoAutoGPT's guide: the cloud deployment ignores these variables

Setting it up

AutoGPT calls this the local transport because Ollama is the usual target, but CHAT_BASE_URL is just a URL — its own guide lists a managed OpenAI-compatible API among the shapes that work. Four variables matter: CHAT_USE_LOCAL=true, CHAT_BASE_URL ending in /v1, CHAT_API_KEY (required — the transport deliberately does not fall back to an OpenAI key), and CHAT_FAST_STANDARD_MODEL as a bare model id. If you leave the advanced, title and simulation models unset, AutoGPT derives them from the standard one so no cloud-only name reaches your endpoint.

AutoGPT Platform v0.8.2, single container — the run's variables with a gateway's values; the advanced-model line is optional
docker run -d --name autogpt \
  --shm-size 2g --ulimit nofile=65536:65536 \
  -p 127.0.0.1:3000:3000 \
  -e AUTOGPT_PUBLIC_URL=http://localhost:3000 \
  -e CHAT_USE_LOCAL=true \
  -e CHAT_BASE_URL=https://api.kunavo.com/v1 \
  -e CHAT_API_KEY=sk-kn-... \
  -e CHAT_FAST_STANDARD_MODEL=claude-haiku-4-5 \
  -e CHAT_FAST_ADVANCED_MODEL=claude-sonnet-5 \
  -v autogpt-data:/data \
  significantgravitas/autogpt:v0.8.2

# Compose installs put the same four CHAT_* lines in autogpt_platform/backend/.env.
# Check them inside the running container:
docker exec autogpt env | grep ^CHAT_

From inside Docker, localhost is the container, not your machine: an endpoint on the host is http://host.docker.internal:<port>/v1 on Docker Desktop, which is how the run reached its stand-in. A hosted endpoint such as Kunavo's needs no special networking. The first boot runs database migrations and takes several minutes; wait for the container to report healthy.

What a turn sends — measured

RequestSizeNotes
POST /v1/chat/completions, first call of a turn41,876 bytesStreaming; 16 tools — among them bash_exec, web_fetch, run_agent, write_workspace_file; body fields messages, model, stream, tools only
Same, after a tool result42,677 bytesFour messages: the tool call and its result appended
TitleAbout 0.6 KBNon-streaming, once per session, same model
Graphiti memoryPer turnPOST /v1/responses with the default Ornith model name, and an embedding call for nomic-embed-text
Context-window probesPer sessionGET /api/ps, /props, /api/v0/models, /v1/models

Sizes were byte-identical across three runs. They are request bytes from a stand-in, not a provider's token count. The request shape is plain OpenAI Chat Completions — no vendor extensions — so an endpoint that serves streaming tool calls is all the chat path needs.

Side traffic to plan for

Everything above went to the same base URL. Two consequences if your endpoint is a chat gateway rather than a full local stack. Embeddings: the marketplace index asked for text-embedding-3-small 200 times at boot, and Graphiti asked for nomic-embed-text each turn; an endpoint without embedding models returns 404, store search falls back to lexical-only, and memory extraction logs tracebacks — the chat itself was unaffected. Kunavo serves no embedding models, so on Kunavo that is the expected state. Memory model names: Graphiti's /v1/responses calls use its own default model name unless GRAPHITI_LLM_MODEL and GRAPHITI_RERANKER_MODEL are set to models your endpoint serves.

What it costs per turn

As rough arithmetic, not a bill: one tool turn sent 84,553 bytes over two calls — about 21,138 input tokens at four characters per token — plus a few hundred output tokens; we assume 600.

Model through KunavoRate per 1M in / outOne tool turn
Claude Haiku 4.5$0.70 / $3.50$0.017
Claude Sonnet 5$1.40 / $7.00$0.034

Conversations grow, so later turns send more. Kunavo bills per token from a prepaid balance with a $10 minimum top-up and no subscription (billing); the catalog amount is a billing floor rather than a cap. Nobody at Kunavo has run AutoGPT against its own endpoint — the run here used a local stand-in — so check the first turn's charge in your usage log.

FAQ

Can AutoGPT use a custom API endpoint?

Yes, for one part of the product and only when you host it. On a self-hosted AutoGPT Platform, the AutoPilot chat agent takes any OpenAI-compatible endpoint through CHAT_USE_LOCAL=true, CHAT_BASE_URL, CHAT_API_KEY and CHAT_FAST_STANDARD_MODEL. We ran v0.8.2's official single-container image that way on October 1, 2026, against a local stand-in endpoint: a chat turn and three tool round trips completed, and the file AutoPilot wrote was read back through the workspace API. The hosted agpt.co service ignores these variables, and the AI Text Generator block inside agent graphs does not read them.

Does the AI Text Generator block use CHAT_BASE_URL?

No. AutoGPT's own guide says the AutoPilot chat path and the block layer read different variables. In v0.8.2's source the block layer's OpenAI and Anthropic clients are built with no base URL, and the only host you can set there is an Ollama host for Ollama's native API — so agent graphs that call the AI Text Generator block keep using the providers AutoGPT ships, whatever CHAT_BASE_URL says.

What is the best or cheapest API for AutoGPT?

For the self-hosted chat path, the cheapest is a model you run yourself — AutoGPT's guide uses Ollama — and the cheapest metered option is whichever OpenAI-compatible endpoint serves a capable tool-calling model at the lowest rate. Each tool turn in our run sent about 85 KB across two chat calls with 16 tool definitions, roughly 21,000 input tokens at four characters per token; at Kunavo's Claude Haiku 4.5 rates that is about $0.017 a turn, and about $0.034 on Claude Sonnet 5. Cheapest listed rate and cheapest finished task are different questions: a model that fumbles tool calls costs more turns.

Why does my AutoGPT endpoint get embedding and /v1/responses requests?

Because other features follow the same base URL. In our run the endpoint received 200 embedding requests for text-embedding-3-small at boot (the marketplace's semantic index), and after every chat turn Graphiti memory sent a /v1/responses call using its default model name and an embedding call for nomic-embed-text. An endpoint without those models answers 404; the chat kept working, and the container logged Graphiti tracebacks. Set GRAPHITI_LLM_MODEL, GRAPHITI_RERANKER_MODEL and GRAPHITI_EMBEDDER_MODEL to models your endpoint serves, or accept degraded memory and lexical-only store search.

How do I check AutoPilot is really using my endpoint?

Three checks, all from AutoGPT's guide and all seen in our run: docker exec into the container and grep the environment for CHAT_; send a chat turn and look for "Using baseline service" in the logs; and confirm your endpoint logged a streaming /v1/chat/completions request with the model you set. Before the first turn, AutoPilot also probes /api/ps, /props, /api/v0/models and /v1/models to learn the context window; endpoints that report none fall back to 32k according to the guide.

Run on October 1, 2026: significantgravitas/autogpt:v0.8.2 (arm64 image, sha256:8322941548c6…) on Docker Desktop, AutoPilot driven through its own chat API with a throwaway local account, against a local recording stand-in that answered one tool call and then text. Three tool round trips, all completed; the written file was read back through the workspace API. Behaviour beyond the run is from AutoGPT's copilot-local-llm guide and the v0.8.2 source. The stand-in is not a provider and not Kunavo.