Home
BlogTalksPapersAbout
Nikita Kozodoi

Nikita Kozodoi, PhD

Senior Applied Scientist at AWS. Designing, engineering and deploying custom AI solutions for customers across industries.

PhD in ML · 8+ years in applied AI/ML · Focused on LLMs and agentic AI

CV
0
Blog Posts
0
Public Talks
0
Papers
0
Paper Citations
0
GitHub Stars
0
Kaggle Medals

Latest work

Latest Blog

AUMOVIO Improves Quality of Automotive Software at Scale Using Multi-Agent AI on Amazon Bedrock

AWS Blog

Defects that clear the compiler, the static analyzer and a code review still surface during vehicle integration testing, the most expensive place to find them. We built a multi-agent detector on Amazon Bedrock AgentCore that returned 30+ net-new high-priority findings in a single scan.

View all 28 blog posts
Latest Talk

Intelligent Document Processing with Generative AI

GITEX AI · Berlin

Live demo of the IDP Accelerator, a scalable, serverless solution for automated document processing and information extraction on AWS. The asset combines generative AI and optical character recognition (OCR), using services such as Amazon Bedrock Data Automation and Amazon Bedrock foundation models to extract, classify, and process documents at scale.

View all 33 talks
Latest Paper

Are We Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

ICML 2026 Workshop on Weight-Space Symmetries

Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematically studying how training duration of domain experts affects the quality of the merged model. We fine-tune experts on five domains (Math, Code, Instruction Following, Multilingual, and Safety) across three model sizes (Qwen 3.5 0.8B, 2B, and 4B), saving checkpoints from 25% to 500% of the optimal training steps and evaluating five merging methods at each duration. Our findings reveal a striking method-dependent pattern: simple averaging degrades sharply with overfitting, while sparsification-based methods achieve their best performance well past the validation optimum. We formalize this through bias-variance decomposition analysis, drawing a parallel to random forests where averaging benefits from high-variance individual learners. These results suggest that training duration and merging method should be chosen jointly rather than independently.

View all 15 papers

Get in touch

Open to collaboration, speaking invitations, and conversations about applied AI.

ContactBuy me a coffee

© 2026 Nikita Kozodoi. All opinions are my own. RSS