Datadog's Test-Driven Production Migrations with Claude + Cursor: A Reusable Methodology
Extending TDD from new code to large-scale legacy refactoring: pin down behavioral contracts with AI first, then use tests as the safety net — Datadog's three-phase engineering practice.


Datadog's Test-Driven Production Migrations with Claude + Cursor: A Reusable Methodology
Extending TDD from new code to large-scale legacy refactoring: pin down behavioral contracts with AI first, then use tests as the safety net — Datadog's three-phase engineering practice.
Large-scale legacy code migrations are something every engineering team dreads: touch one function and a wall of regression tests fails; leave it alone and technical debt keeps piling up. AI tools genuinely speed up code changes, but they speed up "breaking code" just as fast — an AI-driven refactor without a safety net is a disaster.
Datadog's engineering team recently shared a practice worth copying outright: test-driven production migrations using Claude + Cursor. The core idea is to extend TDD (test-driven development) from "writing new code" to "migrating large volumes of legacy code" — use AI to pin down behavioral contracts first, then lean on tests as the safety net.
This tutorial breaks down Datadog's three-phase method and gives you a workflow you can reuse in your own codebase.
Method Overview: Intent → Tests → Migration
Datadog's method has three steps. Each uses AI tooling, but humans stay in the loop at every critical checkpoint:
| Phase | Goal | Tool |
|---|---|---|
| 1. Describe intent | Write down "what this function should do" | Claude |
| 2. Write focused tests | Lock existing behavior into test cases | Human + Claude |
| 3. Execute the migration | Refactor under test protection | Cursor |
The key to this ordering: tests first, then touch the code. It's the exact opposite of the traditional "migrate first, add tests later."
Step 1: Describe Intent with Claude
Where migrations go wrong most often isn't "the code was changed incorrectly" but "the code was changed correctly, yet the behavior shifted" — you assumed some edge case was a bug when it was actually a contract an upstream dependency relied on.
So before touching any code, write down the behavioral intent of every key function:
- What should the function's inputs be?
- What should it return?
- Which edge cases must be preserved?
- What side effects does it have (writing logs, calling external APIs, mutating global state)?
Use Claude to help with this; the prompt looks roughly like:
Below is the source of the function we are about to migrate:
```python
def normalize_event(raw):
...Based on this code, write a "behavioral intent specification" for it:
- What is the function's core responsibility
- Types, ranges, and constraints of the input parameters
- Types, structure, and constraints of the output
- All side effects and edge cases you can identify from the code
- Which behaviors "look like bugs but are actually depended-upon contracts"
Output the specification only; do not modify the code.
> 💡 **Tip**: Getting Claude to focus on point 5 is critical — much of the behavior in legacy code that "looks like a bug" is actually being depended on by some upstream module. Change it and things break.
The output is a natural-language "behavioral contract" that serves as the basis for writing tests in the next step.
## Step 2: Write Focused Tests
With the intent specification in hand, you can write targeted tests. This step is **done jointly by humans and AI**:
- Human: decide which behaviors to lock in (which are core contracts, which are edge cases)
- AI: quickly generate test-case code based on the intent specification
The goal of the tests is not "100% coverage" but **locking in critical behavior** — especially the behavior most likely to break during a migration.
Use Claude to generate a test draft:
```text
Based on the behavioral intent specification below, generate targeted pytest test cases for this function:
[Paste the intent specification from Step 1]
Requirements:
1. Cover all "core contract" behaviors with at least 2 test cases
2. Cover all "edge case" behaviors with at least 1 case each
3. Cover all "looks like a bug but is actually depended upon" behaviors, one case each, with comments
4. Cases must pass both before and after the migration (dependent on behavior, not on implementation)A human reviews the test draft, adds boundary cases the AI didn't think of, and removes redundant cases.
💡 Tip: This step should "nail down" the tests. Once the migration starts, they are the safety net — if any change breaks them, you must stop and confirm whether the tests are outdated or the code has actually gone wrong.
Step 3: Execute the Migration with Cursor
With the tests in place, you can migrate with confidence. This is where Cursor earns its keep — its codebase context capabilities are a great fit for large-scale refactoring.
The working rhythm:
- Migrate in small steps: migrate one module or one group of related functions at a time; no sweeping rewrites
- Run the tests: run the test suite immediately after every small change
- Red-green loop: tests fail → fix code or fix tests → all green → move to the next round
- Human reviews the diff: migration code generated by Cursor must pass human review before it can merge
In Cursor, you can kick off the migration like this:
@Context: select the module to migrate
@Files: reference the relevant test files
Please migrate @normalize_event from the old event format to the new v2 format.
Requirements:
1. Keep all tests in [paste file name] passing
2. Keep every contract described in the behavioral intent specification unchanged
3. Output the changes as a diff, annotating the reasoning behind each change
4. If you hit ambiguous behavior the tests don't cover, stop and ask me — don't decide on your ownThe key design is item 4 — make the AI stop and ask a human when uncertain, rather than deciding on its own. This is the foundation of how test-driven migrations hold onto quality.
Bringing Production Observability into the IDE
Once the migration is done, the biggest risk is "all tests pass, but production breaks" — because some behaviors are beyond what tests can cover, such as issues that depend on real traffic patterns or real data distributions.
This is where Datadog's practice is especially clever: Datadog ships official IDE extensions for Cursor (and VS Code) that bring production observability data (telemetry, logpoints, error tracking) straight into the editor.
If a function misbehaves in production after the migration, you don't need to switch over to the Datadog web app to dig through logs — right inside Cursor you can see the function's real error rate, slow-request samples, and stack traces.
- Datadog Cursor integration docs: https://docs.datadoghq.com/integrations/cursor/
- Datadog IDE extension docs (VS Code/Cursor): https://docs.datadoghq.com/ide_plugins/vscode/
💡 Tip: Even if you're not a Datadog customer, the idea transfers — bring any form of production observability (Sentry, Grafana, a homegrown APM) into your IDE, and let "tests + production data" together form the post-migration safety net.
Verifying the Outcome
For a complete test-driven migration, the verification criteria should be:
- All focused tests pass 100%
- Production metrics (error rate, latency, QPS) show no significant change before vs. after the migration
- Code review leaves no unresolved questions about "did behavior change?"
- The intent specification docs are updated (what is the behavioral contract of the new version?)
FAQ
- What if the legacy code has no tests to begin with?: That's precisely where this method delivers the most value — use AI to help you "fill in" the tests. Deriving tests from a behavioral intent specification is more reliable than deriving them from the code.
- What if the AI-generated tests are themselves wrong?: Human review is always the last line of defense. After the AI generates tests, run the core cases manually at least once to confirm they really are testing the right things.
- Does this method only work at Datadog's scale?: No. The essence of the method is "pin down behavior first, then touch the code," which applies to migrations of any size. Small projects can even skip the Datadog extension — Claude + Cursor + a test suite is enough.
Final Thoughts
The essence of Datadog's practice isn't "they used Claude and Cursor" — it's correctly carrying TDD's mindset into the new territory of large-scale refactoring: use AI to pin down the behavioral contracts first, use tests as the safety net, and only then let the AI change code freely, under protection.
The official case coverage is on InfoQ: https://www.infoq.com/news/2026/07/datadog-ai-production-migration/
Honesty note: This article is based on InfoQ's public coverage of Datadog's engineering practice. The original report gives no precise figures for lines of code migrated or time spent, and this article invents none; it is framed only as a "large-scale, production-grade" engineering case study, not as a new Datadog product.