Download full text
(194.2Kb)
Citation Suggestion
Please use the following Persistent Identifier (PID) to cite this document:
https://nbn-resolving.org/urn:nbn:de:0168-ssoar-66084-2
Exports for your reference manager
Speaker trait characterization in web videos: Uniting speech, language, and facial features
[conference paper]
Abstract
We present a multi-modal approach to speaker characterization using acoustic, visual and linguistic features. Full realism is provided by evaluation on a database of real-life web videos and automatic feature extraction including face and eye detection, and automatic speech recognition. Differen... view more
We present a multi-modal approach to speaker characterization using acoustic, visual and linguistic features. Full realism is provided by evaluation on a database of real-life web videos and automatic feature extraction including face and eye detection, and automatic speech recognition. Different segmentations are evaluated for the audio and video streams, and the statistical relevance of Linguistic Inquiry and Word Count (LIWC) features is confirmed. In the result, late multimodal fusion delivers 73, 92 and 73% average recall in binary age, gender and race classification on unseen test subjects, outperforming the best single modalities for age and race.... view less
Keywords
video; video clip; recording; computational linguistics; Internet; evaluation; social media; experiment; audiovisual media
Classification
Natural Science and Engineering, Applied Sciences
Free Keywords
speaker classification; computational paralinguistics; multi-modal fusion; Linguistic Inquiry and Word Count; LIWC
Collection Title
Proceedings of the 38th International Conference on Acoustics, Speech and Signal Processing (ICASSP 2013)
Conference
38. International Conference on Acoustics, Speech and Signal Processing (ICASSP 2013). Vancouver, 2013
Document language
English
Publication Year
2013
Publisher
IEEE
Page/Pages
p. 3647-3651
DOI
https://doi.org/10.1109/ICASSP.2013.6638338
ISSN
2379-190X
ISBN
978-1-4799-0356-6
Status
Published Version; peer reviewed
Licence
Deposit Licence - No Redistribution, No Modifications