Spam Four Ways: Making Sense of Text Data

Spam Four Ways: Making Sense of Text Data
复制标题

垃圾邮件四种方式:理解文本数据

DOI:
10.1080/09332480.2022.2066414
复制
发表时间:
2022
期刊:
CHANCE
影响因子:
--
通讯作者:
Palmer, Phebe
Palmer, Phebe
中科院分区:
--
文献类型:
--
作者:
Horton, Nicholas J.;Chao, Jie;Finzer, William;Palmer, Phebe

文献摘要

参考文献

被引文献

相似文献

这个世界充满了文本数据,但文本分析在统计学教育中并没有发挥重要作用。我们考虑四种不同的方式,为学生提供机会,探讨电子邮件是否是不必要的信件(垃圾邮件)。主题行中的文本用于识别可用于分类的特征。这些方法包括使用模型激发活动,使用CODAP进行探索,使用专门设计的Shiny应用程序进行建模,以及使用R编写更复杂的分析。这些方法在使用技术和代码方面各不相同,但都有一个共同的目标,即使用数据来做出更好的决策,并评估这些决策的准确性。
The world is full of text data, yet text analytics has not traditionally played a large part in statistics education. We consider four different ways to provide students with opportunities to explore whether email messages are unwanted correspondence (spam). Text from subject lines are used to identify features that can be used in classification. The approaches include use of a Model Eliciting Activity, exploration with CODAP, modeling with a specially designed Shiny app, and coding more sophisticated analyses using R. The approaches vary in their use of technology and code but all share the common goal of using data to make better decisions and assessment of the accuracy of those decisions.
使用统计数据识别垃圾邮件
DOI: --
发表时间: 2015
期刊:
影响因子: --
作者:
D. Nolan;D. Lang
通讯作者: D. Lang
使用 R 进行数据科学
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
Hause Lin
通讯作者: Hause Lin