Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL

·Toolin Editorial Team

Mind Lab open-sources Macaron-V1 (Venti / Coding-Venti / Tall tiers), a Mixture-of-LoRA architecture built on GLM-5.2 744B plus four 1B LoRA experts, with a free hosted API and a UI4A plugin.

Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL

Now that "bigger" is no longer the only direction AI models evolve in, Mind Lab has offered a concrete answer: Macaron-V1 — an open-source coding model built on a GLM-5.2 744B base and assembled from four 1B LoRA experts. Developers can pick it up today: weights fully open on HuggingFace, a free hosted API, and a UI4A visual plugin you can install directly. If you've been looking for an open-source model that "compresses personalized engineering experience into LoRAs as memory," this piece walks you through what it is and how to use it.

What Macaron-V1 Is

Macaron-V1 is a model series open-sourced by Mind Lab. Its core pitch is not stacking parameters higher but treating "experience" as a hot-swappable small module. The whole design is called Mixture-of-LoRA (MoL): one fixed large base (GLM-5.2 744B, with a Qwen3.6 35B variant), plus four LoRA experts at the 1B scale, combined on demand at inference time.

The idea stuffs "personalization / domain knowledge / coding habits" into lightweight LoRAs — what the authors call LoRA-as-memory (experience learning). Instead of training yet another trillion-parameter giant, "experience" exists and updates in the form of small modules. The Macaron naming follows from this: a big base sandwiching several layers of flavored filling.

Model Tiers (Open-Sourced on HuggingFace)

The full weight collection lives at: https://huggingface.co/collections/mindlab-research/macaron-v1

VariantParameter countPurpose
Macaron-V1-Preview-749B754BPreview base, text generation
Macaron-V1-Preview-744B-Merged754BMerged preview base
Macaron-V1-Venti753BGeneral Venti tier, full capabilities
Macaron-V1-Coding-Venti753BCoding-focused Venti, the developer's first pick
Macaron-V1-Tall36BSmall tier, based on the Qwen3.6 35B variant, lightweight and runnable locally

Developers doing coding work should go straight to Coding-Venti; those who want local runs or low latency should pick Tall.

Three Ways to Use It

Route 1: Call the Free Hosted API Directly (Fastest)

No local deployment needed — the fastest way in is the official hosted endpoint:

It is OpenAI-API compatible, so you can plug in with your existing SDK.

Route 2: Deploy the Weights Locally / Self-Hosted

Pull the HuggingFace weights and run them on the standard GLM-5.2 / Qwen3 inference stacks (vLLM, SGLang, and transformers all work). The Venti tier needs inference hardware that can hold 744B; for lighter scenarios, Tall (36B) is enough.

Route 3: Install the UI4A Visual Plugin (Open Source)

UI4A is the Macaron team's visual plugin project, open-sourced at:

git clone https://github.com/MindLab-Research/macaron-artifacts

Repository address: https://github.com/MindLab-Research/macaron-artifacts

UI4A provides a plugin manifest, a WebUI runtime, and GenUI tools, letting Macaron run inside mainstream agent / coding environments like Claude Code, Codex, and Kimi Code — effectively swapping in a "hot-swappable experience" model backend within an existing agent workflow.

Why the MoL Route Deserves Developers' Attention

The mainstream narrative of large-model progress is "more parameters, more capability," but for developers that means two things: deployment keeps getting harder, and the model never "remembers" your project's context. Macaron takes a different path:

  • Stable base, swappable experts: GLM-5.2 744B serves as the stable base, while the four 1B LoRA experts can be trained and replaced independently.
  • Experience exists as modules: your code style, project conventions, and calling conventions can, in principle, be compressed into a LoRA module rather than stuffed into a prompt or a full retrain.
  • Open weights, auditable: all weights are public on HuggingFace, not a black-box API.

This architecture is most attractive for developer teams with their own codebases / knowledge bases — you can imagine training a dedicated LoRA per project and mounting it as needed.

Use Cases

  • A coding copilot for individual developers: connect it to your IDE via the hosted API or the local Tall tier — free and customizable.
  • Team-level code-style / project-convention memory: compress team conventions into LoRA modules so everyone shares the same "experience memory."
  • Agent workflow backend: wire it into Claude Code / Codex / Kimi Code via UI4A as a hot-swappable model layer.
  • MoL architecture experiments: researchers and senior LLM engineers can study LoRA expert composition, experience learning, and modular inference on the open weights.

Final Thoughts

Macaron-V1 is not another open-source model chasing "more parameters, higher benchmarks" — it's an attempt to engineer "experience as hot-swappable memory." What developers can actually hold today: fully open HuggingFace weights (5 tiers), a free hosted API (international and China lines), and the UI4A visual plugin. If you do coding, build agents, or are simply curious about LoRA-expert architectures, run Coding-Venti once — it beats reading ten architecture write-ups.

Official resources: