Skip to main content
POST

Overview

MiniMax-M3 is MiniMax’s latest generation long-context reasoning model, engineered for deep analytical work across massive inputs. With a 1,000,000-token context window and the ability to output up to 131,072 tokens in a single response, M3 is purpose-built for tasks that require sustained logical coherence, multi-step reasoning, and the synthesis of information from heterogeneous data sources. What sets MiniMax-M3 apart is its native support for multimodal input — you can feed it text, images, and video directly into the chat. This makes it an ideal choice for video understanding pipelines, document-heavy analysis workflows, and research scenarios where information lives across multiple media types. The model’s “deep thinking mode” further enhances its ability to break down complex problems into intermediate reasoning steps before producing a final answer, improving accuracy on mathematically intense, code-oriented, or logic-heavy queries. Accessing MiniMax-M3 through Vibetool means you get all of this capability behind a single OpenAI-compatible endpoint, with one API key and one consistent request format. You can swap models, run comparisons, and scale usage without touching your integration logic. Vibetool handles routing, authentication, and billing in one place — so your team stays focused on building rather than managing upstream provider relationships.
Model Slug: minimax/minimax-m3Use this exact slug when making API requests to Vibetool.

Model Specifications

  • Context Window: 1,000,000 tokens (1M tokens)
  • Max Output Tokens: 131,072 tokens (131K tokens)
  • Input Modalities: Text, Image, Video
  • Output Modalities: Text
  • Deep Thinking Mode: Enabled — enhances chain-of-thought reasoning for complex tasks

Pricing

Pricing: See vibetool.ai/pricing for current rates.

Use Cases

  • Long-Document & Video Understanding: Ingest entire research papers, legal briefs, or hours of video content in a single request. M3 can reason across text and visual frames simultaneously without chunking.
  • Deep Analytical Reasoning: Tackle multi-step mathematical proofs, algorithmic design, and complex logic puzzles. The deep thinking mode surfaces intermediate reasoning steps before delivering a final answer.
  • Multimodal RAG & Retrieval Augmentation: Build retrieval-augmented generation pipelines that mix text, images, and video transcripts. M3 can combine information from multiple media types to answer questions that span modalities.
  • Code Generation & Review: Generate, refactor, and explain code across a wide range of programming languages, with the ability to reason over large codebases that fit within the 1M-token context window.

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

application/json
model
enum<string>
required
Available options:
minimax-m3
messages
object[]
required
stream
boolean
default:false
temperature
number
Required range: 0 <= x <= 2
max_tokens
integer

Response

Chat completion

id
string
object
string
Example:

"chat.completion"

model
string
choices
object[]
usage
object