Xiaomi Miloco 2.0: Smart Homes Finally Get a True AI Steward
Xiaomi open sources its whole-home AI solution Xiaomi Miloco 2.0 — multimodal perception, proactive intelligence, and household memory bring the Agent into the smart home ecosystem.


Xiaomi Miloco 2.0: Smart Homes Finally Get a True AI Steward
Xiaomi open sources its whole-home AI solution Xiaomi Miloco 2.0 — multimodal perception, proactive intelligence, and household memory bring the Agent into the smart home ecosystem.
After a long day, you push open the front door late at night and a voice greets you: "Rough day at work — get some rest, and call me if you need anything!" You're pan-frying a steak in the kitchen while scrolling videos, lose track of time, and a voice alert chimes in: "The oil is about to overheat — an estimated 20 seconds until the ideal moment to turn off the heat!" You're away on a business trip, and your child's tablet shuts off automatically when screen time is up and sends you a notification; the AI camera spots that your parents forgot their medication and announces a reminder through the speaker.
This isn't science fiction. Xiaomi has officially released and open sourced Xiaomi Miloco 2.0 — a future-facing open source whole-home AI solution. It is also the first smart home solution in the industry capable of AI proactive service with household memory.
What Miloco 2.0 Is
Miloco 2.0 is the new "AI brain" Xiaomi has fitted onto the smart home for the Agent era. It treats the various Mijia devices as full-modal perception entry points, achieves whole-home understanding through vision, sound, and environmental sensing, and relays user needs to the Agent — closing the loop on AI delivering services in the home.

The key point: all user data stays on-device, with raw data fully isolated from the Agent. Users have complete control of their data, and privacy and security concerns are properly addressed.
Four Core Capability Upgrades
Multimodal Perception: From Single Vision to Layered Sensing
The system processes multiple dimensions at once — visual changes in a space, changes in people, voice and tone, temperature, and more.
Case in point: "proactively remind when water boils unattended" — the camera plus sensors judge that the water has boiled and detect that no one is in the kitchen, then the speaker nearest the user plays a voice alert. The whole flow is natural, efficient, and logical.

Proactive Intelligence: From Rule-Driven to Large-Model Reasoning
Leveraging the strong commonsense reasoning of large models, the system actively observes the state of the user's scene and, based on daily routines and device usage habits, makes its own judgments and offers services proactively.
Case in point: the camera senses the owner has come home, and combined with household memory judges that the arrival time is later than average, inferring the owner has likely been working late — so it proactively offers some comfort.

Continuous Tasks: From One-Shot Execution to Long-Term Tracking
Instead of the traditional "one command, one execution," AI can genuinely stay online around the clock, tracking continuously across multiple time periods.
Case in point: after receiving a birthday reminder instruction, the system proactively orchestrates the home's available devices (various lights, the TV, speakers) to generate a birthday surprise plan, then stays on "standby." When it detects a family member arriving home, it mobilizes the devices to execute the orchestrated plan.
Household Memory: A Dedicated Profile for Every Family Member
When the camera recognizes someone sitting down at the study desk, it recalls that person's household memory and adjusts the lighting to their preference — the husband likes bright, warm light while reading at the computer, while the wife prefers a soft neutral light when writing notes.

Real-World Experience
Strengths
- It remembers, and it knows who you are: no more command-style "one sentence, one execution" interaction — AI genuinely understands each family member's preferences and routines
- Proactive service: without you saying a word, the system judges from the scene state and offers services on its own
- On-device privacy: all data stays on-device, with raw data fully isolated from the Agent
- Open source: available on GitHub, so developers can build their own smart home solutions on top
How It Differs from Traditional Smart Homes
| Traditional Smart Home | Miloco 2.0 |
|---|---|
| Rule-driven, "if... then..." | Large-model reasoning, proactive judgment |
| Single command execution | Continuous task tracking, standing by across time periods |
| No memory, starts from zero every time | Household memory — knows every family member |
| Single modality (voice/app) | Multimodal perception (vision + sound + environment) |
Use Cases
- Elderly living alone: the AI camera monitors missed medication and unusual behavior, proactively alerting family members
- Families with kids: automatic screen-time management — the device powers down on schedule and notifies the parents
- Kitchen safety: monitoring oil temperature, boiling water, and other hazards with timely voice warnings
- Personalized scenes: automatically adjusting lights, temperature, and music to each person's preferences
- Family events: automatic orchestration of birthday surprises and holiday atmospheres
Tip: Miloco 2.0 is an open source solution that needs to be paired with Xiaomi Mijia ecosystem devices. If you're already a Mijia whole-home smart user, this is an experience upgrade; if you're building a smart home, this solution deserves a close look.
If your impression of the smart home still stops at "yelling at a speaker to turn on the lights," Miloco 2.0 represents the next stage — a "J.A.R.V.I.S." with memory that knows who you are and serves you proactively.
Toolin Editorial Team
Categories
Related articles

Xiaomi MiMo-V2.5-Pro-UltraSpeed: A 1T-Parameter Model Running at 1000 tokens/s
Xiaomi's flagship ultra-speed inference model breaks 1000 tokens/s of output on a standard 8-GPU commodity node, with output priced at 18 RMB per million tokens.

Google Ships Two Lightweight Creative Models: Images in 4 Seconds, Video in 10
Google releases the Nano Banana 2 Lite image model and the Gemini Omni Flash video model, built for speed and savings — $0.034 per image and $0.10 per second of video, already embedded in the Gemini App and YouTube Shorts for free.

Claude Tag Turns Claude Into a Team AI Colleague That Lives in Slack
Anthropic releases Claude Tag, evolving Claude Code into a team collaboration agent with shared context, persistent memory, and proactive intervention, now in Beta for Enterprise and Team users.

Doubao Seed 2.1 Pro Hands-On: Coding Crosses the Usability Line, and Multimodal Recognition Holds a Surprise
Volcengine launches the Doubao Seed-2.1 series: the Pro tier reaches production-grade usability for coding and agent calls, photo-based fish recognition outperforms Gemini 3.1 Flash, with 6 tested prompts included.

Headroom: A Token-Slimming Tool for AI Bills, Open-Sourced by a Netflix Engineer, Cuts 90% of Redundant Tokens
Netflix senior engineer Tejas Chopra open-sources Headroom, which losslessly compresses context before it reaches a large model; it has already saved users about $700,000 and 200 billion tokens.

Running a Local LLM Agent with Pi + LM Studio: A Complete Setup Guide
Combine the Pi agent framework with LM Studio's inference server to run Gemma 4 series models on a local Mac for coding, proofreading, and agent tasks, reaching about 75% of frontier-model performance.