Reproducibility Report for ACM SIGMOD 2020 Paper: ``Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks'

Reproducibility Report for ACM SIGMOD 2020 Paper: ``Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks'
复制标题

ACM SIGMOD 2020 论文的再现性报告:“为数据集成任务创建异构关系数据集的嵌入”

DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Saravanan Thirumuruganathan
Saravanan Thirumuruganathan
中科院分区:
--
文献类型:
--
作者:
Riccardo Cappuzzo;Paolo Papotti;Saravanan Thirumuruganathan

文献摘要

被引文献

相似文献

作者提供了简洁且足够清晰的说明,说明在哪里可以找到代码和数据,以及如何重复论文中的主要实验。代码存储库的结构通常是不言自明的,与包含的 README 文件一起,其他研究人员应该能够轻松浏览它。运行代码主要需要在配置文件中设置控制执行的参数。作者提供了带有参数设置的可立即运行的脚本,可以轻松地重现原始论文中最重要的结果。然而,由于原始论文发表后算法和代码的更新,一些与算法质量相关的数字发生了显着变化。尝试重复所提供的脚本未涵盖的实验将需要付出合理的额外努力才能将本文中讨论的设置映射到配置文件中的参数值。
The authors provided concise and sufficiently clear instructions where to find code and data, and how to repeat the main experiments from the paper. The structure of the code repository is generally self-explanatory, and together with the included README file, other researchers should be able to navigate it easily. Running the code mostly requires setting parameters in configuration files that control execution. The authors provided ready-to-run scripts with parameter settings that made it easy to reproduce the most important results from the original paper. However, due to an algorithm and code update after publication of the original paper, a few of the algorithm-quality-related numbers changed significantly. Trying to repeat experiments not covered by the provided scripts would require a reasonable additional effort to map settings discussed in the paper to the parameter values in the configuration files.