Abstract
This demo lets you experience a real-time interactive rap battle between a human and AI or two AI opponents. We develop an end-to-end solution that leverages multiple data modalities and ML models. Our system integrates speech recognition, emotion detection, language modeling, vision, text-to-speech, and voice cloning. The system combines multiple models to process user audio input, generates appropriate and contextually-aware text and synthesizes audio responses aligned with the background beat music. The demo demonstrates the potential of AI systems in creative music applications and offers insights into challenges and opportunities associated with integrating AI systems for interactive entertainment.
Cite
@inproceedings{kozodoi2025towards,
title={Towards {AI} Rapper: Creating an Interactive Rap Battle Experience with Generative {AI}},
author={Kozodoi, N. and Zinovyeva, E. and Afolabi, Z. and Krasheninnikov, E.},
booktitle={NeurIPS 2025 Workshop on {AI} for Music},
year={2025}
}
