Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL
Mind Lab open-sources Macaron-V1 (Venti / Coding-Venti / Tall tiers), a Mixture-of-LoRA architecture built on GLM-5.2 744B plus four 1B LoRA experts, with a free hosted API and a UI4A plugin.


Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL
Mind Lab open-sources Macaron-V1 (Venti / Coding-Venti / Tall tiers), a Mixture-of-LoRA architecture built on GLM-5.2 744B plus four 1B LoRA experts, with a free hosted API and a UI4A plugin.
Now that "bigger" is no longer the only direction AI models evolve in, Mind Lab has offered a concrete answer: Macaron-V1 — an open-source coding model built on a GLM-5.2 744B base and assembled from four 1B LoRA experts. Developers can pick it up today: weights fully open on HuggingFace, a free hosted API, and a UI4A visual plugin you can install directly. If you've been looking for an open-source model that "compresses personalized engineering experience into LoRAs as memory," this piece walks you through what it is and how to use it.
What Macaron-V1 Is
Macaron-V1 is a model series open-sourced by Mind Lab. Its core pitch is not stacking parameters higher but treating "experience" as a hot-swappable small module. The whole design is called Mixture-of-LoRA (MoL): one fixed large base (GLM-5.2 744B, with a Qwen3.6 35B variant), plus four LoRA experts at the 1B scale, combined on demand at inference time.
The idea stuffs "personalization / domain knowledge / coding habits" into lightweight LoRAs — what the authors call LoRA-as-memory (experience learning). Instead of training yet another trillion-parameter giant, "experience" exists and updates in the form of small modules. The Macaron naming follows from this: a big base sandwiching several layers of flavored filling.
Model Tiers (Open-Sourced on HuggingFace)
The full weight collection lives at: https://huggingface.co/collections/mindlab-research/macaron-v1
| Variant | Parameter count | Purpose |
|---|---|---|
| Macaron-V1-Preview-749B | 754B | Preview base, text generation |
| Macaron-V1-Preview-744B-Merged | 754B | Merged preview base |
| Macaron-V1-Venti | 753B | General Venti tier, full capabilities |
| Macaron-V1-Coding-Venti | 753B | Coding-focused Venti, the developer's first pick |
| Macaron-V1-Tall | 36B | Small tier, based on the Qwen3.6 35B variant, lightweight and runnable locally |
Developers doing coding work should go straight to Coding-Venti; those who want local runs or low latency should pick Tall.
Three Ways to Use It
Route 1: Call the Free Hosted API Directly (Fastest)
No local deployment needed — the fastest way in is the official hosted endpoint:
- International: https://mint.macaron.im
- China: https://mintcn.macaron.xin
It is OpenAI-API compatible, so you can plug in with your existing SDK.
Route 2: Deploy the Weights Locally / Self-Hosted
Pull the HuggingFace weights and run them on the standard GLM-5.2 / Qwen3 inference stacks (vLLM, SGLang, and transformers all work). The Venti tier needs inference hardware that can hold 744B; for lighter scenarios, Tall (36B) is enough.
Route 3: Install the UI4A Visual Plugin (Open Source)
UI4A is the Macaron team's visual plugin project, open-sourced at:
git clone https://github.com/MindLab-Research/macaron-artifactsRepository address: https://github.com/MindLab-Research/macaron-artifacts
UI4A provides a plugin manifest, a WebUI runtime, and GenUI tools, letting Macaron run inside mainstream agent / coding environments like Claude Code, Codex, and Kimi Code — effectively swapping in a "hot-swappable experience" model backend within an existing agent workflow.
Why the MoL Route Deserves Developers' Attention
The mainstream narrative of large-model progress is "more parameters, more capability," but for developers that means two things: deployment keeps getting harder, and the model never "remembers" your project's context. Macaron takes a different path:
- Stable base, swappable experts: GLM-5.2 744B serves as the stable base, while the four 1B LoRA experts can be trained and replaced independently.
- Experience exists as modules: your code style, project conventions, and calling conventions can, in principle, be compressed into a LoRA module rather than stuffed into a prompt or a full retrain.
- Open weights, auditable: all weights are public on HuggingFace, not a black-box API.
This architecture is most attractive for developer teams with their own codebases / knowledge bases — you can imagine training a dedicated LoRA per project and mounting it as needed.
Use Cases
- A coding copilot for individual developers: connect it to your IDE via the hosted API or the local Tall tier — free and customizable.
- Team-level code-style / project-convention memory: compress team conventions into LoRA modules so everyone shares the same "experience memory."
- Agent workflow backend: wire it into Claude Code / Codex / Kimi Code via UI4A as a hot-swappable model layer.
- MoL architecture experiments: researchers and senior LLM engineers can study LoRA expert composition, experience learning, and modular inference on the open weights.
Final Thoughts
Macaron-V1 is not another open-source model chasing "more parameters, higher benchmarks" — it's an attempt to engineer "experience as hot-swappable memory." What developers can actually hold today: fully open HuggingFace weights (5 tiers), a free hosted API (international and China lines), and the UI4A visual plugin. If you do coding, build agents, or are simply curious about LoRA-expert architectures, run Coding-Venti once — it beats reading ten architecture write-ups.
Official resources:
- Model collection: https://huggingface.co/collections/mindlab-research/macaron-v1
- UI4A plugin: https://github.com/MindLab-Research/macaron-artifacts
- Hosted API (international): https://mint.macaron.im
- Hosted API (China): https://mintcn.macaron.xin