Abstract
Infrastructure-as-code (IaC) is critical for cloud reliability and scalability, yet LLM capabilities in this domain remain underexplored. Existing benchmarks focus on declarative tools like Terraform and full-code generation. We introduce SWE-InfraBench, a dataset of realistic incremental edits to AWS CDK repositories from real-world codebases. Each task requires modifying existing IaC based on natural language instructions, with correctness verified by passed tests. Results show current LLMs struggle: the best model (Sonnet 3.7) solves 34% of tasks, while reasoning models like DeepSeek R1 reach only 24%.
Cite
@inproceedings{tarasova2025sweinfrabench,
title={{SWE-InfraBench}: Evaluating Language Models on Cloud Infrastructure Code},
author={Tarasova, N. and Balp-Straffon, E. and Iancheruk, A. and Sielskyi, Y. and Kozodoi, N. and Byrne, L. H. and Butler, J. and Jiang, D. and Czelej, M. and Ang, A. and Shah, Y. and Blanco, R. and Ivanov, S.},
booktitle={NeurIPS 2025 Workshop on Evaluating the Evolving {LLM} Lifecycle},
year={2025}
}
