Registration
for the seminar is only possible online via
the central
registration page.
| The goal of the seminar is to autonomously acquire knowledge and critical comprehension of an assigned topic, and present this topic both in writing and verbally. |
Language models assign a probability to a sequence of words, and their estimation has moved from count-based n-gram models over feed-forward and recurrent neural models to the Transformer-based large language models of today. In this topic you should trace this development, work out the central idea and the limitations of each model generation, and describe the current state of the art in greater detail.
Initial Reference:The encoder that maps acoustic features to higher-level representations is a central component of most speech systems, and its architecture has evolved from feed-forward and convolutional networks over (bidirectional) LSTMs to Transformer variants such as the Conformer and, more recently, to state-space models. In this topic you should compare these architectures with respect to their central ideas and to their trade-off between recognition quality, computational cost, parallelization and streaming capability.
Initial Reference:Room 6129
Tel: 0241 80 21630
E-Mail: hwu@ml.rwth-aachen.de