Back to guides
Integration·July 10, 2026·Updated October 4, 2026·8 min read

Kilo Code with the Claude API—base URL configuration (also Cline and Roo Code)

Kilo Code, Cline, and Roo Code all accept a custom OpenAI-compatible endpoint: three fields are enough to run Claude in VS Code, about 30% below Anthropic’s pricing. Tool-by-tool setup, model selection, and the real cost of a session.

Kilo Code, Cline, and Roo Code are the most widely used open-source AI coding agents in VS Code—and all three accept a custom OpenAI-compatible endpoint. That means you can run Claude in your editor with a single sk-kn- key, at about 30 % below Anthropic's official rates for the main models. This guide covers the three-field setup for each tool, which Claude model to choose for each task, and the actual cost of an agentic coding session.

Why these tools require an API key

Kilo Code, Cline, and Roo Code do not include a model: every action (reading a file, suggesting a change, running a command, checking the result) is an API call billed by token. A chat subscription such as Claude Pro does not cover this—you need direct API access. That is exactly what Kunavo provides: pay-as-you-go starting with a $10 top-up, one key for Claude and GPT, a balance that never expires, and card payments from France (VAT and GDPR details are in the Claude API guide for France). Since the agentic loop resends context at every step, the per-token rate is the main lever for a session's cost.

The three fields (all tools)

All three extensions share the same provider model—Kilo Code and Roo Code are forks in the Cline family—so the setup is identical:

parametres-provider
API Provider   OpenAI Compatible
Base URL       https://api.kunavo.com/v1
API Key        sk-kn-...          # à créer sur kunavo.com/app/keys
Model ID       claude-sonnet-5        # ou claude-opus-4-7 / claude-haiku-4-5

Create a key in the dashboard after signing up and adding a $10 top-up—it is shown only once, so store it immediately.

Configure Kilo Code

  1. Install Kilo Code from the VS Code marketplace and open it in the sidebar.
  2. Open the extension settings (gear icon), then the Providers section.
  3. Set API Provider to OpenAI Compatible.
  4. Base URL: https://api.kunavo.com/v1 · API key: your sk-kn-… · model: claude-sonnet-5.
  5. Save, then test with a simple request (“explain this file”) before letting it modify code.

Kilo Code lets you use a different model for each mode (Architect / Code / Debug)—useful for routing planning to a stronger model and execution to a cheaper one.

Configure Cline

  1. Install Cline and open its panel.
  2. Click the settings gear and, under API Provider, choose OpenAI Compatible.
  3. Enter the same three values: base URL https://api.kunavo.com/v1, your key, and model ID claude-sonnet-5.
  4. If Cline asks for model metadata (context window, output limit), use the figures on the model page.

Configure Roo Code

As with the other two: Settings → Providers → API Provider → OpenAI Compatible, then enter the base URL https://api.kunavo.com/v1, key, and model ID. Roo Code's modes (Code / Architect / Ask / Debug) also accept one model per mode.

All three tools also offer a native Anthropic provider with a custom base URL option. Point it to https://api.kunavo.com to use the native Messages API (/v1/messages)—useful because cache_control is passed through as-is and cached input is billed at 10% of the input rate (see the prompt caching docs). Both routes work; the OpenAI-compatible route is the simplest way to get started.

That said, prompt caching itself does not depend on the native route. On a long prompt, Kunavo places the cache breakpoints itself on the OpenAI-compatible route: on the system prompt, on the tool definitions and on the last message once the conversation contains an assistant turn, which in an agentic loop is every step after the first. A cache_control that the extension places there is kept. On the native route, the request is passed on as the extension sent it and no breakpoint is added: what gets cached depends on the breakpoints the extension places.

Which Claude model should you choose?

TaskModelKunavo input / output (per 1M)
Everyday agentic coding (default)claude-sonnet-5$1.40 / $7.00
Previous generation, for prompts already tuned to itclaude-sonnet-4-6$2.10 / $10.50
Difficult refactors and debuggingclaude-opus-4-7$3.50 / $17.50
Small changes, commit messagesclaude-haiku-4-5$0.70 / $3.50

The model field is just a slug on the same endpoint: switching models means changing one word—no new key, no new provider. Full rates are in the Claude API pricing guide and on the pricing page.

The actual cost of a session

Agentic tools consume many tokens by design: each step resends the system prompt, task history, and file context. Realistic figures at Kunavo rates:

UnitTokens (input / output)claude-sonnet-5At Anthropic rates
One agentic step25,000 / 1,200$0.043$0.062
A 20-step task~500k / ~24k~$0.87~$1.24
A heavy day (5 tasks)—~$4.34~$6.20

Run the calculation yourself:

cout_session.py
# Tarifs Claude sur Kunavo (USD par 1M de tokens) : (input, output)
RATES = {
    "claude-haiku-4-5":  (0.70, 3.50),
    "claude-sonnet-5":   (1.40, 7.00),
}

def cout_etape(model, in_tokens, out_tokens):
    i, o = RATES[model]
    return in_tokens / 1_000_000 * i + out_tokens / 1_000_000 * o

# Une étape agentique : l'outil renvoie le contexte + les fichiers lus.
print(cout_etape("claude-sonnet-5", 25_000, 1_200))         # -> $0.0434
# Une tâche réaliste de 20 étapes (éditer, lancer, corriger, répéter) :
print(20 * cout_etape("claude-sonnet-5", 25_000, 1_200))    # -> ~$0.87

How to keep your bill low

  1. Start new tasks instead of stretching the same one. The tool resends the entire conversation at every step: an endless task grows quadratically in cost. New task = short context.
  2. Route by mode. Kilo Code and Roo Code accept one model per mode—Architect on claude-opus-4-7, Code on claude-sonnet-5, quick questions on claude-haiku-4-5.
  3. Use the native Anthropic route for caching. Cached input is billed at 10% of the rate—in a loop that resends a stable prefix at each step, this is the biggest available saving (how it works).
  4. Track spend by key. Give your editor its own key with a spending limit in the dashboard, and check usage to see what a week of agentic coding actually costs.

Frequently asked questions

What base URL should I use for Kilo Code with the Claude API?

Choose the “OpenAI Compatible” provider and set the base URL to https://api.kunavo.com/v1, with a Kunavo key (sk-kn-...) and a Claude model slug such as claude-sonnet-5. The same three fields work in Cline and Roo Code. For the native Anthropic provider, use https://api.kunavo.com (the SDK adds /v1/messages).

Can I use my Claude Pro subscription with Kilo Code?

No—chat subscriptions do not include API access. Kilo Code, Cline, and Roo Code call the API directly and charge by token. On Kunavo, you can top up starting at $10, Claude is served at about 30 % below Anthropic's rates, and your balance never expires.

Which Claude model should I choose for Kilo Code?

claude-sonnet-5 ($1.40/$7.00 per 1M tokens on Kunavo) is the default choice for everyday agentic coding: near-Opus code quality at a lower per-token rate than claude-sonnet-4-6 ($2.10/$10.50) on Kunavo. Move up to claude-opus-4-7 for the toughest refactors. Route simple questions to claude-haiku-4-5. Full rates are in the Claude API pricing guide.

How much does a Claude coding session cost?

Agentic tools resend context at every step: a typical step uses ~25k input tokens / 1.2k output tokens, or about $0.043 on claude-sonnet-5 at Kunavo rates—around $0.87 for a 20-step task, compared with ~$1.24 at Anthropic's official rate. Failed requests are not charged.

What about GDPR compliance?

Kunavo offers neither a DPA (the GDPR Article 28 contract), Zero Data Retention, nor data hosting in the EU: requests are processed in the United States and then transmitted to model providers, whose own retention policies apply. If your project requires a DPA, ZDR, or data residency in the EU, contract directly with Anthropic, OpenAI, or Google under their enterprise terms. Pseudonymizing PII before transmission and other measures on your side are covered in the guide GDPR compliance for LLMs in France.