Efficient Lyrics Extraction from the Web

Efficient Lyrics Extraction from the Web
复制标题

从网络中高效提取歌词

DOI:
--
复制
发表时间:
2006
期刊:
International Society for Music Information Retrieval Conference
影响因子:
--
通讯作者:
J. Korst
J. Korst
中科院分区:
--
文献类型:
--
作者:
G. Geleijnse;J. Korst

文献摘要

被引文献

相似文献

我们提出了一种从网络中提取歌词的新方法。其目的是提取一组歌曲歌词的多个版本。歌词可以通过正则表达式在文本中识别。我们使用文档的投影,通过将其映射到正则表达式来有效地识别文档中的歌词。我们描述了一种方法,通过过滤掉错误的文本(如其他歌曲的歌词)来聚类多个版本的歌词。出于效率的考虑,我们通过比较指纹而不是文本本身来完成这项工作。
We present a novel method to extract lyrics from the Web. The aim is to extract a set of multiple versions of the lyrics to a song. Lyrics can be identified within a text by a regular expression. We use a projection of a document to efficiently identify lyrics within the document by mapping it to a regular expression. We describe a method to cluster the multiple versions of the lyrics by filtering out erroneous texts such as lyrics to other songs. For reasons of efficiency, we do this by comparing fingerprints instead of the texts themselves.