Get API key

kimik2api.comKimi K2 API: First request in five minutes

Kimi K2 API: First request in five minutes

Get your first response in under five minutes using our OpenAI-compatible API. This guide covers the base URL, authentication, and immediate usage patterns for the uncensored model.

Base URL and Authentication

Connect to the kimik2api.com service by pointing your client to https://api.kimik2api.com/v1. You need a valid API key for every request. Generate one on the Get API key page after signing up with just an email and password. No credit card is required to start, and each account gets $0.50 in trial credit valid for 7 days. Keep your key secure; you can regenerate it at any time to revoke the old one.

First Request (cURL)

Test connectivity immediately with a standard chat completion request. Replace YOUR_API_KEY with your actual key. The model ID is uncensored. This endpoint accepts text input and returns text output, supporting up to 64,000 tokens in the combined prompt and completion.

curl https://api.kimik2api.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

If you receive a 401, your key is invalid. A 402 means your prepaid credit is exhausted. A 429 indicates you have exceeded the 300 requests per minute limit.

Python SDK Integration

Use the official openai Python package to interact with the API. Configure the base_url and api_key to route requests to our servers instead of OpenAI's. This allows you to use standard client code without vendor lock-in. The model ID remains uncensored for all completions.

from openai import OpenAI

client = OpenAI(base_url="https://api.kimik2api.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

This setup ensures you can swap the base URL easily. Ensure your environment has the necessary packages installed before running the script.

Node.js SDK Integration

For JavaScript and TypeScript developers, the Node.js OpenAI SDK works identically. Set the baseURL to our endpoint and provide your API key. This approach maintains compatibility with existing OpenAI-compatible client code. You can send messages and receive responses just as you would with other LLM providers.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.kimik2api.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Remember that we do not offer embeddings or image generation, so the standard chat completions endpoint is your primary interface for text-based interactions.

Streaming Responses (SSE)

Enable real-time token generation by setting stream: true in your request body. The API returns Server-Sent Events (SSE) containing partial deltas. This is ideal for chat interfaces where immediate feedback improves user experience. Each chunk contains a portion of the generated text, allowing you to display output as it is created.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Handle the stream events in your client code to update the UI incrementally. This reduces perceived latency for the end user significantly.

Limits, Errors, and Context

Our API supports a 64,000 token context window, covering both the prompt and the completion. The maximum request body size is 8 MB. You are limited to 300 requests per minute per key. If you hit the rate limit, wait before retrying. The model is uncensored but blocks sexual content involving minors. Pricing is transparent: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credit never expires.

Questions and answers

Is this the official Kimi K2 API?

No, we are an independent service hosting an uncensored large language model compatible with the OpenAI API format. We are not affiliated with Moonshot AI or the official Kimi provider. Check their documentation for their specific limits and pricing.

Do I need a credit card for the trial?

No, you only need an email and password to sign up. You receive $0.50 in trial credit valid for 7 days. You can add more credit later using a card or crypto, starting from $10.

What happens if I exceed the rate limit?

You will receive a 429 Too Many Requests error. The limit is 300 requests per minute per API key. You can regenerate your key if needed, but the rate limit applies to the key itself.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API keyRead the docs