SenseNova U1: An Open-Source Infographic Generation Model, 8B Parameters, Runs on a Single GPU
SenseTime's open-source 8B-parameter infographic generation model, Apache 2.0 licensed for commercial use, with stable text rendering and precise layout control at roughly one-tenth the cost of closed-source options.


SenseNova U1: An Open-Source Infographic Generation Model, 8B Parameters, Runs on a Single GPU
SenseTime's open-source 8B-parameter infographic generation model, Apache 2.0 licensed for commercial use, with stable text rendering and precise layout control at roughly one-tenth the cost of closed-source options.
GPT-Image 2 set infographic generation on fire, but it is closed-source, billed by token, at up to $30 per million output tokens. If you need local deployment or secondary development, SenseTime's open-source SenseNova U1 is currently the alternative most worth watching: 8B parameters, Apache 2.0 license, runs on a single GPU, with community-tested costs at roughly one-tenth of the closed-source option.
What SenseNova U1 Is
SenseNova U1 is SenseTime's open-source multimodal generation model family, built on the in-house NEO-unify architecture. It discards the VAE and vision encoder that traditional image models consider indispensable, natively modeling pixels and text in the same representation space. This means the model no longer "translates" images — it thinks in both languages at once, eliminating detail loss and noise from compression at the root.
China's developer community on Hugging Face described it as "achieving pure end-to-end pixel-text modeling."

Core Capabilities of the Infographic-Enhanced Version
On top of the base version, SenseTime further open-sourced SenseNova-U1-8B-MoT-Infographic (the infographic-enhanced version), purpose-tuned for infographic scenarios.
Stable text rendering
The hardest part of infographic generation is text. The SenseNova U1 infographic-enhanced version renders text without visible flaws in dense mixed Chinese-English layouts. In a knowledge-diagram test of "LLM architecture evolution" packed with bar charts, tables, and bilingual Chinese-English text, the years and parameter scales from BERT through GPT-5 were legible at a glance, with no garbled characters.
Precise layout control
In poster-generation tests, the model follows typesetting instructions accurately. Asked for "about 40% of the frame left as whitespace in the center" and "an extremely airy feel," it did not pile on extra decoration but strictly honored the restrained design brief. The pairing of dark serif type with beige paper texture precisely captured the balance between Eastern negative-space poetics and modern layout.
Structured document generation
In an academic-paper page generation test, the model accurately produced a complete arXiv-style page layout with clean formatting; even complex mathematical formulas showed no structural errors, arriving at a directly usable level of finish.

Dense small Chinese text rendering
In a test of densely packed Chinese infographics on enterprise brand operations logic, the model rendered nearly all of the small Chinese text accurately, with a clean, readable layout.

Comparison with GPT-Image 2
The two models have clearly diverged in the infographic generation space:
| Dimension | GPT-Image 2 | SenseNova U1 enhanced |
|---|---|---|
| Design orientation | Visual school: chases the visual impact of light, shadow, and materials | Production-tool school: prioritizes clarity of information structure |
| Text rendering | High quality but font sizes run small | Stable and highly readable |
| Cost | $30 per million output tokens | Roughly 1/10 of the closed-source option |
| Deployment | API calls only | Local deployment, runs on a single GPU |
| License | Closed-source | Apache 2.0, commercial use allowed |
Put simply, GPT-Image 2 suits occasions that chase visual impact, while SenseNova U1 better fits real production scenarios that demand precise information delivery.
Deployment Details
- Model parameters: 8B
- License: Apache 2.0 (commercial use allowed)
- Deployment requirements: a single RTX 5880 (48GB VRAM), with actual VRAM usage around 30GB
- Community support: GGUF quantized weights have been contributed by the community, lowering the deployment barrier further
- Benchmark gains: +6.8 points on BizGenEval (Hard), +18.2 points on IGenBench Q-ACC
Who Should Use It
- Content creators: anyone batch-producing infographics, posters, or knowledge diagrams
- Dev teams: teams that need locally deployed infographic generation and want to avoid ongoing API fees
- Academic/office scenarios: anyone generating structured documents or presentation pages
- Commercial users: the Apache 2.0 license allows commercial use, with nothing to worry about