Cancer information on the world wide Web: gross characteristics.

Cancer information on the world wide Web: gross characteristics.
复制标题

DOI:
10.1093/jnci.djh048
复制
发表时间:
2004-02
期刊:
Journal of the National Cancer Institute
影响因子:
--
通讯作者:
Craig W. Trumbo
Craig W. Trumbo
中科院分区:
其他
文献类型:
--
作者:
Craig W. Trumbo

文献摘要

被引文献

相似文献

基于网络的癌症信息已被检查(1-7),但其总体特征尚未被描述。这项研究提供了2001年3月和2003年3月的站点描述。在2001年3月获取的Web内容样本中,使用迭代试验来确定最简单的搜索字符串,该字符串将产生没有错误材料的结果。使用的搜索字符串是“癌症或肿瘤或肿瘤*或肿瘤或恶性* -天文* -星座-回归线-蟹-生肖。”搜索引擎是Northern Light (www.northernlight.com)和Alta Vista (www.altavista.com)。在2001年,搜索引擎谷歌不允许布尔搜索。这两个搜索引擎的前200个结果被合并在一起,任何剩余的错误链接和重复链接都被删除,包括学术期刊和门户网站。结果是306个通用资源定位器(url)。2003年3月采集了第二次样本。当时,北极光已经暂停运营,谷歌引入了布尔搜索。第二个样本是专门用谷歌做的,再次取前400个退货并删除相同种类的材料。结果是326个url。这些样本代表了当时一个偶然的搜索者可以获得的材料主体的一个公平的画面。编码器被训练来检查一些明显变量的内容:主页状态,页面货币的日期指示的提供,以及15种癌症(加上“其他”)信息的存在。内容宽度的指数被计算为页面上存在的癌症形式数量的总和。为了衡量网页在网络中的嵌入程度,通过使用http://www.linkstoyou.com/Checklinks.htm网站来评估指向每个给定页面的链接数量。该变量报告指向AltaVista中找到的页面的链接的数量,不包括内部链接。第一个示例中的页面寿命是通过检查URL在1年和2年后(2002年3月和2003年3月)是否仍然有效来评估的。最后,计算两个可读性分数。Flesch Reading Ease评分为1-100分,分数越高越容易阅读(最好是60-70分)。Flesch-Kincaid年级水平量表将文本置于美国小学水平。两者都是通过Microsoft Word中提供的实用程序运行的。为了执行可读性分数,从页面中随机选择一段文本(例如,一段文字)。当打开的页面没有提供足够的文本时,随机选择一个链接并跟随到下一页。没有必要再继续下去了。当我们分析页面的特征时,可以在两个样本之间观察到一些有趣的差异(表1)。2003年的样本包含了更大比例的主页,链接更彻底的页面,以及可读性更好的内容。然而,必须指出的是,即使是在2003年样本中发现的10年级水平的可读性也比推荐给普通读者的难度要大。当我们分析癌症内容的特征时,唯一没有增加代表性的癌症类别是儿童癌症和其他。在两个样本中,各种癌症的排名只有轻微的变化(Spearman秩序相关性,rs = 0.70; P = 0.003)。正如广度得分的跳跃所反映的那样,2003年样本中的典型网站似乎比2001年样本中涉及更多种类的癌症。这可能代表着从更专业的网站向服务更广泛受众的网站的转变。或者,这可能是整体站点开发和扩展的结果。搜索引擎的选择标准也在不断发展,以支持更广泛定义的网站。对页面寿命的评估发现,1年后,79%(95%置信区间[CI] = 73%至78%)的url仍然存活,2年后,58% (95% CI = 52%至63%)的url仍然存活。在两年的研究期间,主页比网页更有可能保持活跃。1年后,96%的主页存活,73%的网页存活(χ2 =; P<.001); 2年后,94%的主页存活,47%的网页存活(χ2 <48; P<.001)。从最广泛的意义上说,我们的分析表明,在这项研究的两年里,网络上有关癌症的内容在很多方面都有所改善。几乎所有形式的癌症在更大比例的主页上都有更好的表现,这些主页通过网络进行了更广泛的链接。可读性也有所提高。这种评估Web内容的方法可能会带来一些希望。迄今为止,所有关于癌症的网络研究都集中在单一癌症上作为范例,将结果推广到更广泛的内容。在这项研究中看到的不同癌症之间的差异可能表明,这种概括并不完全可靠。也就是说,必须指出的是,对网络癌症内容的检查所描绘的画面充其量只是部分的。对Web内容的广泛类别进行抽样是有问题的,而且可用材料的绝对数量几乎不可能对其进行完整的描述。搜索引擎的变化,尤其是跨时间跨度的变化,使得比较变得困难。虽然有人可能会说,网站的数量是不可知的,所有可以观察到的都是由特殊的和不断发展的搜索引擎提供的结果。从这个意义上说,我们的研究并不是在描述网络,而是在描述搜索结果的本质。无论如何,这种Web内容评估方式的改进可以为那些致力于向公众提供癌症信息的人提供有价值的反馈。
Web-based cancer information has been examined (1–7), but its gross characteristics have not been previously described. This study provides such a description of sites in March 2001 and in March 2003. In the sample of Web content taken in March 2001, iterative trials were used to determine the simplest search string that would yield results free of erroneous material. The search string used was “cancer OR oncology OR neoplas* OR tumor OR malignan* -astro* -horoscope -tropic -crab -zodiac.” Searches were executed with the search engines Northern Light (www.northernlight.com) and Alta Vista (www.altavista.com). In 2001, the search engine Google did not allow for Boolean searches. The top 200 results from the two search engines were combined, and any remaining erroneous links and duplicates were removed, including sites for academic journals and Web portals. The result was 306 universal resource locators (URLs). A second sample was taken in March 2003. At that time, Northern Light had suspended operation and Google had introduced Boolean searches. The second sample was done exclusively with Google, again taking the top 400 returns and removing the same variety of material. The result was 326 URLs. These samples represent a fair picture of the body of material available to a casual searcher at those times. Coders were trained to examine the content on a number of manifest variables: home page status, provision of a dated indicator of page currency, and the presence of information on 15 forms of cancer (plus “other”). An index of content breadth was calculated as the sum of the number of forms of cancer present on the page. To gauge the embeddedness of the page in the Web, the number of links pointing to each given page was assessed by use of the site http://www.linkstoyou.com/Checklinks.htm. The variable reports the number of links pointing toward the page found in AltaVista, excluding internal links. The longevity of the pages in the first sample was assessed by checking to see if the URL was still live 1 and 2 years later (March 2002 and March 2003). Finally, two readability scores were calculated. The Flesch Reading Ease score provides a rating of 1–100, with higher scores being easier to read (a score of 60–70 is desirable). The Flesch–Kincaid Grade Level scale places the text on a U.S. grade school level. Both were run through the utility provided in Microsoft Word. To execute the readability scores, a block of text (e.g., a paragraph) was randomly selected from the page. When the opening page did not provide sufficient text, a link was randomly selected and followed to the next page. It was not necessary to proceed further. When we analyzed the characteristics of the page, several interesting differences were observed between the two samples (Table 1). The 2003 sample included a much greater percentage of home pages, of pages that were more thoroughly linked to, and of content with improved readability. However, it must be noted that readability at even the 10th grade level that was found in the 2003 sample is more difficult than is recommended for a general audience. Table 1 Page and content characteristics When we analyzed the characteristics of the content on cancer, the only categories of cancer that did not increase in their representation were childhood cancers and other. There was only a slight change in the rankings of the various cancers across the two samples (Spearman rank order correlation, rs = .70; P =.003). As reflected by the jump in the breadth score, it appears that the typical Web site in the 2003 sample addresses a greater variety of cancers than that in the 2001 sample. This might represent a shift away from more specialized Web sites and toward sites that serve a broader audience. Alternatively, this is possibly a consequence of overall site development and expansion. Search engine selection criteria may have also evolved to favor more broadly defined sites. The assessment of page longevity found that 1 year later 79% (95% confidence interval [CI] = 73% to 78%) of URLs were alive and that 2 years later 58% (95% CI = 52% to 63%) were still alive. Home pages were much more likely than pages to remain alive across the 2-year study period. At 1 year, 96% of home pages were alive versus 73% of pages (χ2 =; P<.001), and at 2 years, 94% of home pages remained alive versus 47% of pages (χ2 <48; P<.001). In the broadest sense, it might be argued that our analysis shows that in a number of ways the Web’s cancer content may have improved during the 2 years of this study. Nearly all forms of cancer are better represented on a greater percentage of home pages that are more extensively linked through the Web. Readability has also improved. This approach to evaluating Web content may hold some promise. All of the studies done to date on the Web’s representation of cancer have focused on single cancers as exemplars, generalizing the results to the broader content. The differences seen in this study between specific cancers may suggest that such generalizations are not entirely reliable. That said, it must also be pointed out that the picture painted by this examination of the Web’s cancer content is still only partial at best. Sampling broad categories of Web content is problematic, and the sheer volume of material available makes a complete characterization nearly impossible. Changes in search engines, especially across longer time spans, make comparisons difficult. Although it might be argued that the population of Web sites is unknowable and all that can ever be observed are the results provided by idiosyncratic and evolving search engines. In that sense, our study is not characterizing the Web but, rather, is characterizing the nature of search returns. In any case, the refinement of this manner of Web content evaluation could provide valuable feedback for those working to provide cancer information to the public.