Registration
for the seminar is only possible online via
the central
registration page.
| The goal of the seminar is to autonomously acquire knowledge and critical comprehension of an assigned topic, and present this topic both in writing and verbally. |
Speech large language models extend text-based large language models to spoken input and output, either by representing speech as a sequence of discrete tokens or by coupling a speech encoder to a pre-trained language model through an adapter. In this topic you should classify the existing approaches along these lines, describe selected ones in greater detail, and discuss what they gain over classical dedicated systems and at which cost.
Initial Reference:State-of-the-art recognition quality comes with model sizes that are expensive to train and, above all, to deploy, which is addressed by efficient architectures, by compression techniques such as knowledge distillation, pruning and quantization, and by faster search. In this topic you should give a structured overview of these techniques, pay attention to how efficiency is actually measured, and compare the reported trade-offs between recognition quality and cost.
Initial Reference:Room 6129
Tel: 0241 80 21630
E-Mail: hwu@ml.rwth-aachen.de