Topic Analysis project in CS5604, Spring 2016: Extracting Topics from Tweets and Webpages for IDEAL

Topic Analysis project in CS5604, Spring 2016: Extracting Topics from Tweets and Webpages for IDEAL
复制标题

CS5604 中的主题分析项目,2016 年春季:从 IDEAL 的推文和网页中提取主题

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
R. K. Vinayagam
R. K. Vinayagam
中科院分区:
--
文献类型:
--
作者:
Sneha Mehta;R. K. Vinayagam

文献摘要

被引文献

相似文献

本次提交包括项目报告、最终演示文稿、LDA代码、测试数据集及其结果。在压缩文件夹"www.example.com"中,我们包含了用于处理推文的LDA Scala源代码(lda_v1.scala)和用于网页分析的JavaScript文件。压缩文件夹"TopicAnalysis-TestData & Results.zip"包含来自奥巴马医改收集的清理过的推文集合和网页。在同一个文件夹中,我们还包含了每个集合的主题结果和一个用于解释集合ID的PDF文件。
This submission includes the project report, final presentation, LDA code, test datasets and its results. In the compressed folder, "TopicAnalysis-code.zip", we have included the LDA Scala source code (lda_v1.scala) for processing Tweets and a JAR file for web page analysis. The compressed folder, "TopicAnalysis-TestData&Results.zip" contains cleaned Tweet collections and web pages from the Obamacare collection. In the same folder, we have also included the topic results for each collection and a PDF file to interpret the collection IDs.