Multi-label Classification using Logistic Regression Models for NTCIR-7 Patent Mining Task

Multi-label Classification using Logistic Regression Models for NTCIR-7 Patent Mining Task
复制标题

DOI:
--
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
Akinori Fujino;Hideki Isozaki
Akinori Fujino;Hideki Isozaki
中科院分区:
其他
文献类型:
--
作者:
Akinori Fujino;Hideki Isozaki

文献摘要

相似文献

我们设计了一个基于机器学习方法的多标签分类系统的NTCIR-7专利挖掘任务。在我们的系统中,我们采用了逻辑回归模型的每一个国际专利分类(IPC)代码,确定IPC代码分配的研究论文。使用任务组织者提供的专利文献训练logistic回归模型。为了减轻逻辑回归模型对专利文献的过拟合,我们利用研究论文集,采用特征加权和成分选择方法设计专利文献的特征向量。使用NTCIR 7专利挖掘任务的日本子任务的测试集,我们证实了我们的多标签分类系统的有效性。
We design a multi-label classification system based on a machine learning approach for the NTCIR-7 Patent Mining Task. In our system, we employ a logistic regression model for each International Patent Classification (IPC) code that determines the IPC code assignment of research papers. The logistic regressionmodels are trainedby usingpatentdocuments providedby task organizers. To mitigate the overfitting of the logistic regression models to the patent documents, we design the feature vectors of the patent documents with feature weighting and component selection methods utilizing a research paper set. Using a test collection for the Japanese subtask of the NTCIR7 Patent Mining Task, we confirmed the effectiveness of our multi-label classification system.