Use NeverMetered with OpenAI SDK

The official Python and TypeScript clients, with Chat Completions or the Responses API.

OpenAI ChatOpenAI Responses
Model
qwen3.8-27b
8,192 max output tokens on the free planReplace YOUR_API_KEY with your key.
  1. 1

    Create an API key

    Create a key in API keys and name it after the tool, for example “OpenAI SDK”. A dedicated key per tool lets you see its usage separately and give it its own parallel-stream cap. The key is shown once; the snippets below use YOUR_API_KEY where it goes.

  2. 2

    Install the SDK

    pip install openai
  3. 3

    Export your key

    export NEVERMETERED_API_KEY="YOUR_API_KEY"
  4. 4

    Stream a chat completion

    import os
    from openai import OpenAI
    
    client = OpenAI(base_url="https://api.nevermetered.com/v1", api_key=os.environ["NEVERMETERED_API_KEY"])
    
    stream = client.chat.completions.create(
        model="qwen3.8-27b",
        messages=[{"role": "user", "content": "Reply with the word ready"}],
        stream=True,
    )
    for chunk in stream:
        if chunk.choices:
            print(chunk.choices[0].delta.content or "", end="", flush=True)
  5. 5

    Or use the Responses API

    Chain turns with previous_response_id; follow-ups are routed to the engine that holds the conversation, which keeps its cache warm.

    first = client.responses.create(model="qwen3.8-27b", input="Pick a number between 1 and 10.")
    second = client.responses.create(
        model="qwen3.8-27b",
        previous_response_id=first.id,
        input="Now double it.",
    )
    print(second.output_text)
  6. 6

    Verify it works

    Run the script; it prints the model's reply. The request shows up in Requests within a few seconds, with its input, cached and output tokens.

Tips
  • Cap your client's concurrency (for example with a semaphore) at your parallel-stream limit plus queue allowance to avoid 429s from your own traffic. Your plan allows up to 8,192 output tokens per request.
  • Keep system prompts and tool definitions byte-for-byte stable between requests to get cached input tokens.

Troubleshooting

404 model_not_found
Use a model id exactly as GET /v1/models lists it, for example qwen3.8-27b. Plans can restrict which models a key may call.
429 concurrency_limit_exceeded
All your parallel streams are busy and your queue allowance is full. Lower the tool's parallelism (subagents, batch size), or get more streams on your plan.
Answers stop mid-sentence
The response hit your plan's output cap of 8,192 tokens (finish reason length or max_tokens). Ask for shorter steps, or upgrade for a higher cap.
403 email_verification_required
Free requests need a verified email. Click the link in your verification email (you can resend it from the dashboard), or upgrade.
429 daily_quota_exceeded
You've used today's free requests; agents send one request per step, so they go quickly. The allowance resets at 00:00 UTC, and paid plans have no daily limit.
503 / 504 request_queue_timeout
The service stayed busy longer than the queue timeout and nothing was generated. Retry with backoff, and give the tool a generous request timeout.