Five Models Tested: Is Qwen3.7 Max Actually Good at Coding

·Toolin Editorial Team

Qwen3.7 Max climbed to second place globally on competitive coding leaderboards, behind only Claude Opus 4.7. Using four tasks — liquid simulation, hexagonal 2048, a metro museum, and a browser operating system — this hands-on test compares the coding performance of Qwen3.7 Max, GPT-5.5, Gemini 3.5 Flash, DeepSeek V4, and Claude Opus 4.7.

Five Models Tested: Is Qwen3.7 Max Actually Good at Coding

Alibaba's latest flagship model Qwen3.7 Max has taken second place on competitive coding leaderboards, ahead of GPT-5.5, Gemini 3.5 Flash, and DeepSeek V4 Pro, and behind only Claude Opus 4.7. On traditional coding benchmarks like Terminal Bench and SWE Bench, it is also the top domestic model.

Leaderboard numbers look nice, but how does it hold up in practice? We tested Vibe Coding capability with five models, four tasks, and one-sentence prompts.

Getting the Model and Pricing

Qwen3.7 Max is currently available on the Alibaba Cloud Bailian platform, with 1 million free tokens for new users. Limited-time half-price: 6 yuan/million tokens input, 18 yuan/million tokens output.

Compared with Opus 4.7 and GPT-5.5, Qwen3.7 Max has a clear price advantage; against DeepSeek's low prices, though, it is still considerably more expensive.

Test 1: Liquid Simulation Animation

Prompt: "Build an animation with HTML+CSS+JS that simulates liquid sloshing inside a container; dragging the container changes the tilt angle."

Qwen3.7 Max: Completed the task smoothly and even threw in extras like color customization, shaking, and liquid volume adjustment. Solid performance.

DeepSeek V4: Fairly simple, but no errors.

GPT-5.5: The liquid effect is a bit odd, and the wave animation breaks the illusion.

Gemini 3.5 Flash: The bottle hides behind the control panel and has to be dragged out manually. But it offers the most customization options.

Claude Opus 4.7: The bottle is too crude, and under violent motion the liquid sloshing looks like an audio waveform bouncing.

Test 2: Hexagonal 2048

Prompt: "Build a playable 2048, but with hexagonal tiles."

This test examines whether a model can understand game logic on a non-standard grid.

Qwen3.7 Max: Good-looking page, playable, but numbers occasionally stack in the wrong position.

DeepSeek V4: The board is hexagonal, yet keyboard control only offers the four WASD directions.

Claude Opus 4.7: The best performer. It genuinely understood the honeycomb rules — tile movement directions fully match hexagonal logic.

GPT-5.5: Leveraging Codex, after generating it can open a browser preview by itself and scrape console information to fix the code. But its mouse direction control is worse than Opus 4.7's.

Gemini 3.5 Flash: Piled on a large amount of extra content — three background themes plus built-in 8-bit space sound effects — maxing out the experience.

Test 3: Metro Museum Website

Prompt: "Design a theme website named Metro Museum; it should feel strongly immersive."

The intent was to see models present subway information and logos from different cities, with an artistic overall style.

Qwen3.7 Max: Vertical text laid out like subway trains, but the overall feel is chaotic.

Gemini 3.5 Flash: The standout. It built metro-themed merchandise and a custom commemorative ticket-stub generator — enter a name, pick a station, and a vintage-style commemorative ride ticket is generated in real time.

GPT-5.5: Nice page style, but far too little content — it missed that a metro museum should be an information-display website.

DeepSeek V4: It designed ticketing souvenirs and a driving experience, but the final deliverable did not present those features.

Test 4: Browser Operating System

Prompt: "Build a complete browser operating system in HTML."

Qwen3.7 Max: Added a nice desktop landscape image as a bonus, but the whole thing leans simple.

Gemini 3.5 Flash and GPT-5.5: Tied for best. Both designed the entire OS in detail, with a distinct style and complete interactions.

DeepSeek V4: The simplest.

Hooking Qwen3.7 Max into Codex

In testing, generating on the Qwen website alone produced worse results than going through an Agent product like Codex. Here is how to hook it up:

Configuration Steps

  1. Get an API Key from the Alibaba Cloud Bailian platform
  2. Edit the ~/.codex/config.toml config file and add the model information
  3. Also update your computer's environment variables (.bash_profile or .zshrc) with the API Key
  4. Run codex in the terminal to start; the model switches from GPT-5.5 to Custom
# Environment variable example (add to .zshrc or .bash_profile)
export QWEN_API_KEY="your-api-key-here"

The same method also works for hooking models like DeepSeek, MiniMax, and Kimi into Codex.

Codex configuration

Note: After hooking it up you may hit the error stream disconnected before completion: InternalError.Algo.InvalidParameter: The "function.arguments" parameter of the code model must be in JSON format. This is because Bailian's streaming output format is not the standard OpenAI protocol, so Agent tool calls are less stable. When this happens, all you can do is wait for an official fix or open a new session.

Even Better with a Skill on Top

After installing a front-end design Skill in Codex (https://github.com/Leonxlnx/taste-skill), the same one-sentence prompt makes Codex automatically invoke the design Skill to nail down design positioning and concepting, and the final output is considerably better than generating directly on the Qwen website.

Results with the Skill

Verdict

One-line summary: Qwen3.7 Max has genuinely improved a lot at coding, but in one-sentence Vibe Coding scenarios it still cannot consistently beat GPT-5.5 and Gemini 3.5 Flash.

But hooked into an Agent product like Codex and paired with Skills, the results take a qualitative leap. It also confirms a trend: raw model capability is no longer enough — "surrounding capabilities" like memory, orchestration, verification, and reasoning sustainability matter just as much.

Related articles

Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises
AI Products

Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises

A hands-on review of ByteDance's Doubao Seed 2.1 Pro: agent coding and multimodal capability have crossed the production-ready line, including rebuilding front-end interactions from screenshots, at a price nearly 80% lower than Claude Opus 4.6.

Toolin Editorial Team
Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation
AI Products

Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation

Deemos has released Hyper3D Rodin Gen-2.5, the first to bring an LLM-like Thinking mechanism to 3D generation — million-polygon models in 4 seconds, 10-million-polygon precision, and native 12K texturing.

Toolin Editorial Team
Hands-On with WeChat's "Xiaowei" AI Assistant: 12 Entry Points Covering Chat, Content, and Documents
AI Products

Hands-On with WeChat's "Xiaowei" AI Assistant: 12 Entry Points Covering Chat, Content, and Documents

WeChat's native AI assistant Xiaowei is in gray testing. Its main model is the in-house WeLM; it can search chat history, summarize official-account articles, and invoke local-life services, with a second confirmation required for sensitive operations.

Toolin Editorial Team
DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
AI Products

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories

The Gaoling School of AI at Renmin University of China has released DeNovoSWE, the first long-horizon training set for generating complete repositories from documents, with 4818 real task instances; Qwen3-30B improved from 5.8% to 47.2% on BeyondSWE-Doc2Repo.

Toolin Editorial Team
Hands-On with Baidu's DuMate: A Tutorial for the Homegrown Office Agent, from Installation to Automation
AI Tutorials

Hands-On with Baidu's DuMate: A Tutorial for the Homegrown Office Agent, from Installation to Automation

A step-by-step walkthrough of Baidu DuMate's request → authorize → execute → deliver pipeline — installation, interface, hands-on office scenarios, and automation, all covered in one article

Toolin Editorial Team
Claude Tag: Making AI a True Member of Your Team
AI Products

Claude Tag: Making AI a True Member of Your Team

Anthropic has launched Claude Tag, embedding Claude into Slack workflows with shared context, persistent memory, and proactive intervention — a new paradigm for team collaboration.

Toolin Editorial Team