Home
BlogTalksPapersAbout
Nikita Kozodoi

Talks

33Talks
Nikita Kozodoi

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code

Home
Talks
SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code
Conference1 delivery

Infrastructure-as-code (IaC) is critical for cloud reliability and scalability, yet LLM capabilities in this domain remain underexplored. Existing benchmarks focus on declarative tools like Terraform and full-code generation. We introduce SWE-InfraBench, a dataset of realistic incremental edits to AWS CDK repositories from real-world codebases. Each task requires modifying existing IaC based on natural language instructions, with correctness verified by passed tests. Results show current LLMs struggle: the best model (Sonnet 3.7) solves 34% of tasks, while reasoning models like DeepSeek R1 reach only 24%.

PosterPaper

Where it was given

  • NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle · San Diego · 2025

Related content

Paper

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code

2025
Talk

Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection

2025
Blog

Are We Merging the Right Models?

2026
← All talks

© 2026 Nikita Kozodoi. All opinions are my own. RSS