Home
WorkBlogTalksPapersAbout
Nikita Kozodoi

Talks

35Appearances
Nikita Kozodoi

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code

  1. Home
  2. Talks
  3. SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code
Conference2 deliveries

Infrastructure-as-code (IaC) is critical for cloud reliability and scalability, yet LLM capabilities in this domain remain underexplored. Existing benchmarks focus on declarative tools like Terraform and full-code generation. We introduce SWE-InfraBench, a dataset of realistic incremental edits to AWS CDK repositories from real-world codebases. Each task requires modifying existing IaC based on natural language instructions, with correctness verified by passed tests. Results show current LLMs struggle: the best model (Sonnet 3.7) solves 34% of tasks, while reasoning models like DeepSeek R1 reach only 24%.

PosterPaper

Where it was given

  • Agentic AI Summit · Berkeley · 2026
  • NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle · San Diego · 2025

Related content

Paper

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code

2025
Blog

SWE-InfraBench + Kiro: What Happens When a Coding Agent Tackles IaC?

2026
Talk

Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection

2025
← All talks

Pages

  • Work
  • Blog
  • Talks
  • Papers
  • About

Social

  • LinkedIn
  • GitHub
  • Google Scholar
  • X / Twitter
  • Instagram

Contact

  • n.kozodoi@icloud.com
  • Buy me a coffee
  • Download CV
  • RSS feed
  • Berlin, Germany

© 2026 Nikita Kozodoi. All opinions are my own.