Hyra-1.0: Tencent Hunyuan's Scientific Discovery Agent

·Toolin Editorial Team

A research agent capable of recursive self-improvement, covering AI training optimization, open math problems, quantum computing, and drug design, with multiple records broken.

Hyra-1.0: Tencent Hunyuan's Scientific Discovery Agent

Tencent Hunyuan has released its first research agent, Hyra-1.0 (Hunyuan Research Agent). It is not a chatbot, and not yet another coding model — it is an agent that can, like a researcher, "propose a hypothesis → run experiments → distill lessons → propose new approaches based on those lessons," achieving recursive self-improvement (RSI). It has already produced value in more than a dozen real environments — AI model training optimization, GPU kernel optimization, open math problems, qubit routing, drug design — and has refreshed existing records on multiple public tasks. This piece walks you through its design philosophy, measured results, and open-source locations.

What Hyra Is

Hyra's (Hunyuan Research Agent) goal is explicit: convert intelligence and compute into real, usable gains with a simple, general, performance-oriented framework. It follows The Bitter Lesson put forward by Rich Sutton — what wins in the long run is not ever more elaborate handcrafted rules, but giving AI a bigger search space so compute can keep paying off.

At Tencent, Hyra does not just run public benchmarks; it can learn and evolve inside product systems, real AI R&D pipelines, and natural science and industrial scenarios.

Harness Design: Two Core Roles

Hyra's framework (the harness) is deliberately kept lightweight. At its core is one loop: explore and propose better solutions based on accumulated experience, looping until the agent ends on its own or the budget runs out, then return the best solution in history. It consists of two roles:

1. Context Agent

  • Maintains an experience bank (Experience Bank, EB).
  • Writes each solution proposed during exploration, plus its evaluation results, into the EB.
  • Distills the EB into "inspirations" — bundles of context that may include code, files, artifacts, and logs.
  • Feeds inspirations into the task queue.

2. Proposal Agent

  • Picks up one inspiration context from the task queue each time.
  • Based on reflection over that context, writes a new solution (a solution/ folder with a solve.sh entry point).
  • Runs and scores it in a completely fresh sandbox.

The whole thing is an asynchronous producer-consumer pipeline: the Context Agent keeps distilling inspirations and filling the task queue; the Proposal Agent drains the queue; finished solutions are committed back to the EB. A semaphore governs the pipeline, keeping resources healthily saturated.

What If There Is No Evaluator: The Two-Layer Loop

For problems shipped without an evaluator, Hyra upgrades the single-layer loop into a two-layer loop:

  1. Hyra first designs an initial evaluator based on the task description.
  2. The inner loop repeatedly optimizes the solution against that evaluator.
  3. When the inner loop ends, Hyra uses the experience history in the EB to further improve the evaluator (raising efficiency, reducing reward hacking, strengthening evaluation granularity).
  4. The upgraded evaluator opens the next round of the loop.

Take an example: for the task "design a world-champion-level chess AI," the initial evaluator might be a ladder of several random AIs scored by Elo rating; as the loop advances, the opponents escalate step by step into stronger AIs from the EB, achieving co-evolution of eval and solution.

💡 Tip: behind this mechanism sits an insight the Tencent team stresses repeatedly — automated research cannot only optimize the final score; the eval itself must become part of the research loop. In experiments Hyra did encounter solutions that abnormally depressed the metric by "converting a causal language model into approximate bidirectional attention to leak future tokens" — exactly the kind of cheating only an eval upgrade can catch.

Measured Results (Experiment Artifacts Open-Sourced)

All artifacts mentioned below are open-sourced in the GitHub repo: https://github.com/Tencent-Hunyuan/hyra-results

AI for AI: Optimization on the R&D Pipeline

On three tasks from Recursive's automated AI research system:

  • NanoChat Autoresearch: jointly trading off model architecture, optimizer, learning-rate strategy, data flow, and execution efficiency within a fixed budget, ultimately driving Validation BPB down to 0.9015.
  • NanoGPT Speedrun: on a task long optimized by the community with very little headroom left, further cut the time to reach 3.28 validation loss to 76.4s.
  • SOL-ExecBench: jointly optimized 235 kernels against real workloads, reaching a Mean SOL of 0.771.

All three results surpass the baselines reported by Recursive.

AI for Science: From Math to Drug Design

  • Open math problems: gathered 55 open math problems from databases such as EinsteinArena and Erich's Packing Center, many of which had seen no progress for decades. Hyra refreshed the best-known results on 29 of them.
  • Astronomical data: given monthly sunspot counts from 1749-1932, Hyra discovered a recurrence formula that achieved R²=0.77 prediction accuracy on out-of-sample records spanning nearly a century (1932-2026). It can also discover scaling laws from AI training logs.
  • Neural network compression: designed a Transformer with only 15 trainable parameters that performs 10-digit addition — a 58.3% reduction from the 36 parameters of the public record on Adderboard.
  • Quantum computing: on IBM Q20, routing efficiency improved 44.4% versus the classical SABRE algorithm, with CX-equivalent two-qubit gates down from 69 to 24 (a 65.2% reduction).
  • Drug design: given the PARP1 target pocket, Hyra's candidate molecule scored significantly above the approved PARP inhibitor olaparib (-9.77) on the combined binding-druggability score, reaching -10.60. The team stresses these are preliminary simulation screening results, still requiring stricter simulation, synthesis, and wet-lab validation.

AI for Fun: Matches, Modeling, Music

  • Othello bot: through continuous self-play optimization, evolving from MinMax into a hybrid combining AlphaZero-style PUCT with MCTS, it ranks 3rd among the 730 competing programs on Botzone.
  • 3D modeling: designed an object's three-dimensional structure from a single 2D reference image and generated a renderable 3D model — closer to the reference image than models produced by Claude Code's goal mode.
  • Music composition: given a melody, wrote complete harmonies for multiple instruments; over 7 rounds of evolution it developed from a conservative hymn-style harmony into full multi-instrument orchestration with diminished and sixth chords.

Use Cases and the Road Ahead

Tencent states plainly that Hyra's current demos are still very preliminary; the next step is defining more "valuable tasks that can be continuously improved." Beyond public benchmarks, Tencent's product systems, real AI R&D pipelines, and natural science and industrial scenarios will all be decomposed into executable, verifiable problems with clear feedback, and fed to Hyra.

More worth watching is this evolutionary closed loop: better scaffolds produce better data and experience, forging stronger models; stronger models in turn drive stronger scaffolds and more discoveries. The new generation of Hy models, and even parts of Tencent's product line, will keep co-evolving with Hyra.

Research agents may well become one of the defining forms of the next AI generation — AI is moving from "helping humans finish work" further toward "helping humans create knowledge."

Who Should Follow This

  • AI researchers and engineers: those following recursive self-improvement and automated AI research.
  • Scientific computing and cross-disciplinary practitioners: open math problems, quantum computing, drug design, astronomical data analysis.
  • Product and technical decision-makers tracking agent evolution: watch how AI moves from benchmark tests into real research pipelines.