watch Minimax M3 1 Flash · Research

MiniMax M3.1 Flash: 1M-Context Coding Model That Always Thinks

Data graphic: MiniMax M3.1 Flash, a multimodal coding model that always thinks before it answers, shown with a huge "1M" for its 1-million-token context window. Two smaller stats give its 5 effort levels low to max, default max and a 66.25% score 53/80 points on KingBench 3, an independent 8-task test.
AK

Threat intelligence editor · Updated Oct 2, 2026, 1:04 PM EDT

MiniMax's M3.1 Flash preview brings a 1M-token context, image and video input and five effort levels to MiniMax Code. What it does, and how it tests.

MiniMax has launched MiniMax-M3.1-Flash-Preview, a coding model with a 1-million-token context window. It went live on 27 September 2026 inside MiniMax Code, the company's agentic coding tool. MiniMax describes it as a "frontier multimodal coding model" built for agentic reasoning, tool use, coding and long-context work. It is the first M3.1 model to ship.

What M3.1 Flash is

According to MiniMax's API documentation, M3.1 Flash:

  • Holds 1,000,000 tokens of context, which MiniMax pitches at long documents, whole codebases and multi-step agent sessions.
  • Takes text, image and video input, so screenshots, diagrams and screen recordings can go into the same prompt as the code.
  • Always thinks before it answers. Reasoning cannot be switched off: a request that sends thinking: {"type": "disabled"} or effort: "none" is rejected with HTTP 400 and the message that the model "requires adaptive thinking".
  • Has five effort levels: low, medium, high, xhigh and max. Higher levels think longer, produce more output tokens and take more time. If a request leaves effort out, the model runs at max.

That last default matters in practice. A tool that does not set effort gets the slowest, most token-hungry setting on every call. For quick edits and autocomplete-style tasks, set low or medium explicitly.

Data graphic: MiniMax M3.1 Flash always thinks. A request passes through one of five effort levels low, medium, high, xhigh, max into adaptive thinking that can't be turned off, then to the answer. Leaving effort unset runs it at max, and disabling thinking or sending effort none returns an HTTP 400 error.

Every M3.1 Flash request runs through always-on adaptive thinking: leaving effort unset means max, and turning thinking off returns HTTP 400.

How it performs

MiniMax has not published a single benchmark score for M3.1 Flash, nor a model card, parameter count, speed figure or open weights.

The clearest independent result so far comes from the AICodeKing channel's KingBench 3, a set of eight build-it-from-scratch coding tasks scored out of 10 each. Tested on 28 September, M3.1 Flash scored 53 of 80 (66.25%). The best result in the same run was Claude Opus 5.5 at 93.75%.

The scores were uneven:

  • Strong on 3D geometry. It scored 9/10 on a folding-table task, with smooth animated 3D geometry.
  • Weak on interactive simulations. An elevator simulation crashed on load because the code called a position value as a function (3/10). An archery game drew its targets in the wrong place and paused its timer between shots (4/10).

Some sites are also circulating figures such as a 73.8% SWE-bench Verified score, 165 tokens per second and a $0.10 per million input-token price. MiniMax has not published any of these, and we could not trace them to a primary source. Treat them as unverified until MiniMax releases its own numbers.

How to get it

M3.1 Flash is "available only through M Plan and MiniMax Code for now". That means two things:

  • No standalone pay-as-you-go API. You can't call it by the token from your own stack yet.
  • No third-party routers. You can't put it behind a gateway alongside the models you already use.

To try it:

  1. Open MiniMax Code, on the desktop app (macOS or Windows) or the web.
  2. Pick MiniMax-M3.1-Flash-Preview as the model.
  3. Set the effort level per task, rather than leaving it on the default max.

From 1 to 7 October 2026, MiniMax is giving subscribers unlimited use of M3.1 Flash in MiniMax Code. That makes this week a good time to test it on your own repositories.

Data graphic: In AICodeKing's independent KingBench 3 test, MiniMax M3.1 Flash scored 9/10 on a 3D folding-table task but only 4/10 on an archery game and 3/10 on an elevator simulation. Its total was 53/80 66.25% across all 8 tasks, an average of 6.6/10 per task.

M3.1 Flash scored 9/10 on a 3D geometry task but 4/10 and 3/10 on two interactive builds, for 53/80 (66.25%) overall on AICodeKing's independent KingBench 3.

Should you use it?

M3.1 Flash is worth testing now on big-context work, such as reading a large repository or reviewing long agent sessions, and on visual tasks that need image or video input. Two limits apply until MiniMax publishes benchmarks and a per-token price:

  • Keep it in a side-by-side trial, not as the default model in a production pipeline.
  • Check its output on anything interactive or stateful, where the independent tests show it is weakest.

Sources

Keep reading

All latest →
  1. watchResearchGemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5: Is Google's New Model Actually Better?6 min
  2. stableResearchLing 3.1 Flash vs GLM 5.3 Flash vs Qwen 3.8 Flash Next: Three Bets on Cheap Agent Models6 min
  3. elevatedResearchLing-3.1-flash Is Ant Group's Best Flash Model Yet, and It Scores 87.9 on CyberGym6 min
  4. elevatedResearchChatGPT Pro 500: What $500 a Month Buys6 min
  5. stableResearchClaude Sonnet 5.5: Same Price as Sonnet 5, More Work per Dollar — and How to Run It Efficiently7 min
  6. watchResearchBest Local LLM for 16GB in 2026: RTX 5060 Ti, 4060 Ti and Mac mini M4 (Coding First)13 min