XtraGPT: An AI Paper Revision Tool Built on Full-Paper Context
An ACL 2026 paper that uses 20 academic writing criteria and full-paper context modeling to turn AI paper revision from generic polishing into controlled, targeted edits


XtraGPT: An AI Paper Revision Tool Built on Full-Paper Context
An ACL 2026 paper that uses 20 academic writing criteria and full-paper context modeling to turn AI paper revision from generic polishing into controlled, targeted edits
Drop a paper paragraph into ChatGPT and say "make it better," and what comes back is usually just smoother sentences — while the academic problems that actually matter, like missing argument chains, inconsistent terminology, and an underpowered motivation, go completely unaddressed. XtraGPT, proposed by Professor Bingsheng He's team at the National University of Singapore (accepted to ACL 2026), does exactly one thing: controllable paper revision based on full-paper context. It's for researchers and students with a draft in hand who need to refine it to academic standards.
What Is XtraGPT
XtraGPT is not an "AI writes your paper for you" tool. It is a revision-only collaborative system:
- The author must have a complete paper draft
- The author selects the passage to revise and gives revision instructions
- The model produces targeted revision suggestions based on the context of the entire paper
- The author reviews the diff and decides whether to accept it

The core difference from ChatGPT: ChatGPT only sees the paragraph you paste in, while XtraGPT sees the full paper's 16,384-token context.
Core Features
Full-Paper Context Awareness
This is XtraGPT's most crucial capability. When revising a motivation paragraph, the model simultaneously considers the problem definition in the introduction, the assumptions in the methods section, and the results in the experiments section, ensuring the revised paragraph stays consistent with the whole paper.

Ablation studies show that removing the full-paper context drops performance by about 15 points, while removing criteria grounding costs only about 5. Context matters more than training strategy.
20 Academic Writing Criteria
XtraGPT compiled 20 section-level criteria covering six parts of a paper:

These criteria come from writing guides, review rubrics, and expert revision experience. You don't need to pick criteria manually — the model already learned, during training, to map your natural-language instructions to the corresponding academic revision strategies.
Controllable Revision
You express the revision intent in natural language, and the model carries out targeted edits. For example:
- "Strengthen this paragraph's contribution statement"
- "Make the methods description more rigorous"
- "Make the experimental analysis more convincing"
The model makes changes targeted at the specific problems you flag, all while preserving consistency with the rest of the paper.
Hands-On Results
Paper-Level Validation
The team took 54 ICLR 2024 papers, revised them paragraph by paragraph with XtraGPT, then had them scored by an AI-Scientist judge:
| Dimension | Improvement |
|---|---|
| Contribution | +7.9% |
| Presentation | +12.5% |
| Soundness | +6.4% |
| Overall rating | 6.08 -> 6.73 (+0.65) |

Undetectable by AI Detectors
Across 7000 test samples, outputs from both XtraGPT-7B and XtraGPT-14B were classified by Fast-DetectGPT and Binoculars as landing on the human-written side.
Resources
- Paper: https://arxiv.org/pdf/2505.11336
- Code: https://github.com/Xtra-Computing/XtraGPT
- Model (14B): https://huggingface.co/Xtra-Computing/XtraGPT-14B
- PaperDebugger: https://arxiv.org/abs/2512.02589
The models come in multiple sizes from 1.5B to 14B, built on the Qwen-2.5 and Phi families, and can be deployed and run locally.
Use Cases
- Graduate students polishing before submission: once advisor feedback comes back, use XtraGPT to revise paragraph by paragraph against the review criteria
- Non-native English researchers: revising with full-paper context preserves terminological and logical consistency far better than polishing passages in isolation
- Iterating on drafts: you already have the experiments and the ideas and need to lift the draft from "readable" to "submittable"
How It Differs from General-Purpose LLMs
| Dimension | ChatGPT / Claude | XtraGPT |
|---|---|---|
| Context | Only the passage you paste in | The full 16k-token paper |
| Revision type | Generic polishing | Targeted revision against academic criteria |
| Author control | The model decides what to change | The author specifies what to change and how |
| Training data | General conversation | 140,000 real paper revision pairs |
| Output | Replaces the original text | Diff format, reviewable and rejectable |