Skip to content

    Model family

    Qwen on FlexAI.Every variant. One key

    Qwen is Alibaba's open model line on FlexAI. 5 variants run serverless on the OpenAI-compatible API, with 18 more available as dedicated endpoints, spanning chat, code, multimodal, embeddings. One API key serves every variant.

    Variants

    Every served variant in the family, with live serverless pricing.

    Serverless · pay per token

    ModelContextPriceStatus
    Qwen3.6-35B-A3B256K$0.1 / $0.15 per M Serving
    Qwen3 Coder 30B A3B256K$0.07 / $0.26 per M Serving
    Qwen3 30B A3B Thinking 2507256K$0.08 / $0.28 per M Serving
    Qwen3.5 9B250K$0.1 / $0.15 per M Serving
    Qwen3 8B40K$0.02 / $0.1 per M Serving

    Dedicated endpoints · reserved GPUs

    ModelContextPriceStatus
    Qwen3 Coder 480B A35B256KDedicatedDedicated
    Qwen3.5 397B A17B256KDedicatedDedicated
    Qwen3 235B A22B 2507256KDedicatedDedicated
    Qwen3-VL 235B A22B Instruct256KDedicatedDedicated
    Qwen3 Coder Next256KDedicatedDedicated
    Qwen3-Next 80B A3B Instruct256KDedicatedDedicated
    Qwen-AgentWorld-35B-A3B256KDedicatedDedicated
    Qwen2.5 Coder 32B Instruct128KDedicatedDedicated
    QwQ 32B128KDedicatedDedicated
    Qwen2.5 32B Instruct128KDedicatedDedicated
    Qwen3-VL 32B Instruct256KDedicatedDedicated
    Qwen3-VL 30B A3B Instruct256KDedicatedDedicated
    Qwen3-30B-A3B-Instruct-2507256KDedicatedDedicated
    Qwen3.6-27B256KDedicatedDedicated
    Qwen3-Embedding-8B40KDedicatedDedicated
    Qwen3-VL 8B Instruct256KDedicatedDedicated
    Qwen3.5-4B256KDedicatedDedicated
    Qwen3-Embedding-4B40KDedicatedDedicated

    Which variant for what

    Pick by the role you're filling. Same key for all of them.

    Flagship

    Qwen3.6-35B-A3B

    Qwen3.6-35B-A3B is the largest served serverless variant. Reach for it first.

    Fast & economical

    Qwen3 8B

    Qwen3 8B is the smallest serverless variant. Lowest latency and cost.

    Qwen3 Coder 30B A3B runs generate in coding agents

    Call the flagship

    OpenAI-compatible. Swap the model id for any variant above.

    curl https://tokens.flex.ai/v1/chat/completions \
      -H "Authorization: Bearer $FLEXAI_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "Qwen3.6-35B-A3B-FP8",
        "messages": [{"role": "user", "content": "Hello from FlexAI"}]
      }'

    Run Qwen on one API key

    Every Qwen variant, serverless and dedicated, behind one OpenAI-compatible key.

    $10/month in free credits for your first 3 months