Home
BlogTalksPapersAbout
Nikita Kozodoi

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code

Home
Papers
SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code
ConferenceJune 2025

Tarasova, N., et al.

NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle

IEEE Summit 2026

PDFPublisher

Abstract

Infrastructure-as-code (IaC) is critical for cloud reliability and scalability, yet LLM capabilities in this domain remain underexplored. Existing benchmarks focus on declarative tools like Terraform and full-code generation. We introduce SWE-InfraBench, a dataset of realistic incremental edits to AWS CDK repositories from real-world codebases. Each task requires modifying existing IaC based on natural language instructions, with correctness verified by passed tests. Results show current LLMs struggle: the best model (Sonnet 3.7) solves 34% of tasks, while reasoning models like DeepSeek R1 reach only 24%.

Cite

@inproceedings{tarasova2025sweinfrabench,
  title={{SWE-InfraBench}: Evaluating Language Models on Cloud Infrastructure Code},
  author={Tarasova, N. and others},
  booktitle={NeurIPS 2025 Workshop on Evaluating the Evolving {LLM} Lifecycle},
  year={2025}
}

Related content

Talk

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code

2025
Paper

Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection

2025
Talk

Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection

2025
← All papers

© 2026 Nikita Kozodoi. All opinions are my own. RSS