Overview
In this workshop, we aim to shine a spotlight on new approaches in audio-visual generation, as well as covering a wide range of topics related to audio-visual learning, where the convergence of auditory and visual signals unlocks new opportunities for advancing AI-driven creativity, understanding, and perception. We hope our workshop can bring together researchers, practitioners, and enthusiasts from diverse disciplines in both academia and industry to delve into the latest developments, challenges, and breakthroughs in audio-visual generation and learning. Some of the main topics that will be covered are:
- Audio-Visual Generation: sounding video generation, audio-driven talking head animation, video-to-audio generation, etc.
- Audio-Visual Foundation Models: large-scale audio-visual pre-training, foundational video/audio generation and understanding models, etc.
- Real-World Applications: deepfake detection, speech recognition, speech enhancement, video dubbing, audio-visual localization, etc.
- Ethics and Safety: navigating the implications of synthetic media and increasingly capable machine perception.
Call for Papers
Submissions CLOSED
We invite submissions related to audio-visual generation and learning, or any of the topics above. Submissions are now closed. Thank you to everyone who submitted!
We invite submissions for work that has already been accepted in other conferences/journals (e.g., CVPR, ICCV, ECCV, ICML, ICLR, NeurIPS, ICASSP, Interspeech, IJCV, TPAMI, TMLR, etc.), as well as new, work-in-progress research efforts/ideas.
Authors should use the official ECCV 2026 author kit listed at
https://eccv.ecva.net/Conferences/2026/SubmissionPolicies.
Please ensure that you use the template in camera-ready format with all authors and affiliations
included.
Accepted papers will be posted on this website as-is. In the Overleaf template, you can search for the
comments saying
"TODO FINAL" in main.tex to find what needs to be updated.
The submitted papers must be at most 4 pages, excluding references. Note that we also welcome shorter (e.g., 2-page) extended abstracts.
NOTE: If your paper has already been accepted to ECCV 2026, you may submit the full camera-ready paper as is — there is no need to create a new shortened version.
Accepted presenters will have their paper posted on our website, will present a short 5-minute talk at the workshop, and may also present a poster during our poster session. Since they are extended abstracts, they will not be included in the main conference proceedings.
Important Dates
| Submission Deadline | July 15, 2026 |
| Notification of Acceptance | August 1, 2026 |
| Workshop Date | September 9, 2026 |
Submit Your Paper
Submissions are now closed. Thank you for your interest!
Keynote Speakers
Industry Talks
Organizers
Program
Schedule ECCV 2026 — Sep 9th · Room: Malmömässan K1 · All times in CEST
| 09:00 – 09:10 | Welcome & Opening Remarks |
| 09:10 – 10:00 | Oral Paper Session 1 ORAL |
| 10:00 – 10:40 | Keynote Speaker 1 — Prof. Dima Damen KEYNOTE |
| 10:40 – 11:00 | Morning Coffee Break BREAK |
| 11:00 – 11:30 | Industry Talk 1 — Dr. Nikita Drobyshev (Cantina) INDUSTRY |
| 11:30 – 12:10 | Keynote Speaker 2 — Prof. Ruohan Gao KEYNOTE |
| 12:10 – 13:00 | Lunch Break BREAK |
| 13:00 – 14:00 | Poster Session POSTER |
| 14:00 – 14:40 | Keynote Speaker 3 — Prof. Joon Son Chung KEYNOTE |
| 14:40 – 15:10 | Industry Talk 2 — Prof. Ioannis Patras (Queen Mary University of London, Tavus) INDUSTRY |
| 15:10 – 15:50 | Keynote Speaker 4 — Prof. Kristen Grauman KEYNOTE |
| 15:50 – 16:20 | Industry Talk 3 — Google DeepMind INDUSTRY |
| 16:20 – 17:00 | Oral Paper Session 2 ORAL |
| 17:00 – 17:30 | Keynote Speaker 5 — Prof. Tae-Hyun Oh ORGANIZER TALK |
| 17:30 – 17:35 | Closing Remarks & Wrap-Up |
Accepted Papers
Oral Paper Session 1 09:10 – 10:00
-
Presenter: Nathaniel Cohen
-
Presenter: TBA
-
Presenter: Yaofeng Su
-
Presenter: Junseok Ahn
-
Presenter: Junyoung Seo
Oral Paper Session 2 16:20 – 17:00
-
Presenter: Viacheslav Vasilev
-
Presenter: Jisoo Park
-
Presenters: Kai Hsu Tsai, Yong Wei Fu, Hung I Yang
-
Music-Driven Dance Video Generation via PCA-Whitened Latent Diffusion with Rotary Position EmbeddingPresenter: Nokap Tony Park
All accepted papers will also present a poster during the Poster Session (13:00 – 14:00).