Back to blog
Implementation guide·May 25, 2026·8 min read

Build a Japanese RAG chatbot with Claude — a 5,000-document knowledge base in 30 lines

A complete RAG implementation that makes 5,000 internal documents searchable with Claude Sonnet 4.6. About 0.9 yen per query after prompt caching. Embeddings are not provided by Kunavo; only that step is billed directly by OpenAI. Includes Japanese-specific token usage, hallucination countermeasures, and a production-readiness checklist.

A complete example of building a RAG chatbot that makes your internal knowledge base searchable with Claude, using a Japanese-language corpus of 5,000 documents. We’re publishing it with about 30 lines of code and an estimate of monthly operating costs.

Overview

  1. Embed documents as vectors (text-embedding-3-large)
  2. Embed user questions with the same model and search for the top 5 matches
  3. Send the retrieved context to Claude Sonnet 4.6 to generate an answer
  4. Use Prompt caching to reduce repeat-call costs by 90%

1) Embed the documents (once only)

embed.py
# 1) 5,000 件の社内文書を埋め込みベクトル化
from openai import OpenAI

# 埋め込みは Kunavo では提供していません(/v1/embeddings は実装済みの
# ワイヤフォーマットですが、有効なモデルがないため呼ぶと失敗します)。
# 埋め込みは OpenAI に直接、生成は Kunavo に — クライアントは 2 つ持ちます。
embedder = OpenAI(api_key="sk-...")                     # OpenAI 本家
kunavo = OpenAI(
    api_key="sk-kn-...",
    base_url="https://api.kunavo.com/v1",
)

documents = load_documents()  # [{id, text, metadata}, ...]

# バッチ 100 件ずつで埋め込みを呼ぶ
batches = [documents[i:i+100] for i in range(0, len(documents), 100)]
all_vectors = []
for batch in batches:
    resp = embedder.embeddings.create(
        model="text-embedding-3-large",
        input=[d["text"] for d in batch],
    )
    all_vectors.extend(resp.data)

# Postgres + pgvector に保存(または Pinecone, Qdrant など)
save_to_vector_db(documents, all_vectors)

At 5,000 documents × 500 tokens on average, that’s 2.5 million tokens. Kunavo does not provide embeddings, so this step alone is billed directly by OpenAI—— check OpenAI’s pricing page for the rate. We don’t put a number here because a price for something we don’t sell can become outdated without anyone noticing. Run this process once and save the results to pgvector.

2) Search and answer generation at query time

query.py
# 2) ユーザー質問 → 関連文書を検索 → Claude に投げて回答
def answer(question: str) -> str:
    # 質問を埋め込み化
    q_embed = embedder.embeddings.create(   # 埋め込みは OpenAI 直
        model="text-embedding-3-large",
        input=[question],
    ).data[0].embedding

    # 上位 5 件を検索
    relevant = vector_search(q_embed, top_k=5)

    # Claude に context + question を渡す
    context = "\n\n---\n\n".join(d["text"] for d in relevant)
    resp = kunavo.chat.completions.create(  # 生成は Kunavo
        model="claude-sonnet-4-6",
        messages=[
            {
                "role": "system",
                "content": (
                    "あなたは社内ナレッジ ベースをもとに正確に答えるアシスタントです。"
                    "提供された context にない情報については「情報がありません」と答えてください。"
                ),
            },
            {"role": "user", "content": f"# Context\n{context}\n\n# 質問\n{question}"},
        ],
        max_tokens=800,
    )
    return resp.choices[0].message.content

Tokens used per query: question embedding (~$0.000005) + about 5K input tokens of Claude Sonnet 4.6 (context) × $1.20/1M = $0.006 + about 500 output tokens × $6.00/1M = $0.003. About $0.009 per query (about ¥1.4).

3) Reduce costs further with Prompt caching

Anthropic prompt caching works because the same system prompt is sent every time:

cache.py
# 3) Cache_control で再呼び出しコストを 90% 削減
# 同じ context(社内ドキュメント)を何度も使うので、Anthropic prompt caching が効く

resp = client.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[
        {
            "role": "system",
            "content": [{
                "type": "text",
                "text": SYSTEM_PROMPT_AND_GUIDELINES,  # 安定したシステムプロンプト
                "cache_control": {"type": "ephemeral"},
            }],
        },
        {"role": "user", "content": f"# Context\n{context}\n\n# 質問\n{question}"},
    ],
)
# 2回目以降の同じシステムプロンプトはキャッシュヒット → 入力料金の 10%

If the system prompt remains identical at 3,000 tokens, that portion of the input is billed at 10% of the price from the second call onward. Measurements show that the cost drops to about $0.004 per query (about ¥0.6).

Estimated monthly operating costs

UsageCostNotes
Initial embedding$0.25One-time
1,000 queries/dayAbout $270/monthWithout caching
1,000 queries/dayAbout $110/monthWith caching
10,000 queries/dayAbout $1,100/monthWith caching, Slack bot scale

Japanese-specific considerations

  • Token usage: Japanese can use 2-3 times as many tokens as English (for encoding-related reasons involving kanji and kana). Reduce the context from 5K to 3K and rely on search accuracy
  • Embedding model: text-embedding-3-large supports multiple languages and performs well enough in Japanese. If you have many internal terms, we recommend maintaining a separate synonym dictionary
  • Claude vs Gemini 2.5 Pro: Claude is more precise with long summaries and citations; Gemini 2.5 Pro offers broader search. The standard approach is to start with Sonnet 4.6 and escalate to Opus 4.7 as needed
  • Preventing hallucinations: Explicitly instruct the model in the system prompt to “answer ‘I don’t have that information’ if it isn’t in the context.” If that still lets hallucinations through, limit max_tokens and have the model return source document IDs so the UI can display citation links

Production checklist

  • Embedding index update process (daily or weekly rebuild via cron)
  • Fallback for failures (two-stage setup such as Claude → Gemini 2.5 Flash)
  • Rate limits and spending caps (e.g., a $50 daily cap per key at /app/keys)
  • Logging and monitoring (track daily costs in the usage dashboard)
  • Act on Specified Commercial Transactions (for customers in Japan) — /legal/tokutei-shoutorihiki

To get started, sign up for free. Top up from $10 and begin pay-as-you-go usage. For details, see /docs/quickstart or /docs/caching.