面向中文网络百科的属性和属性值抽取 Attribute and Attribute Value Extracted from Chinese Online Encyclopedia期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

按检索

面向中文网络百科的属性和属性值抽取

引用本文：	贾真,杨宇飞,何大可,刘胜久,尹红风.面向中文网络百科的属性和属性值抽取[J].北京大学学报(自然科学版),2014,50(1):41.

作者姓名：	贾真杨宇飞何大可刘胜久尹红风

作者单位：	西南交通大学信息科学与技术学院, 成都610031;

基金项目：	国家自然科学基金(61170111,61202043,61262058);中国科学院自动化研究所复杂系统管理与控制重点实验室开放课题(20110102);中央高校基本科研业务费专项基金(SWJTU11ZT08)资助

摘要：	针对面向中文网络百科条目文章的属性和属性值抽取, 提出一种无监督方法。此方法将属性值看做命名实体, 利用频繁模式挖掘和关联分析, 从文本中抽取类别属性; 采用自扩展方法为属性建立触发词表; 基于属性触发词和属性值实体标注挖掘属性值抽取模式, 利用层次聚类算法获取高质量的模式。在互动百科中采集的数据集上进行实验, 结果表明所提方法行之有效。
关键词：	知识获取属性抽取非结构化文本模式挖掘
收稿时间：	2013-06-20
Attribute and Attribute Value Extracted from Chinese Online Encyclopedia

JIA Zhen,YANG Yufei,HE Dake,LIU Shengjiu,YIN Hongfeng.Attribute and Attribute Value Extracted from Chinese Online Encyclopedia[J].Acta Scientiarum Naturalium Universitatis Pekinensis,2014,50(1):41.

Authors:	JIA Zhen YANG Yufei HE Dake LIU Shengjiu YIN Hongfeng

Institution:	School of Information and Science Technology, Southwest Jiaotong University, Chengdu 610031;

Abstract:	An unsupervised approach is proposed to extract attribute and attribute value from Chinese online encyclopedia entry articles. Attribute values are viewed as named entities and class attributes are extracted based on frequent patterns mining and association analysis. A bootstrapping method is used to find attribute trigger words for each attribute. Attribute value extraction patterns are generated automatically from sentences which contain attribute trigger words and named entity tags of attribute value. Hierarchy clustering algorithm is applied to obtain reliable patterns. Experimental dataset are collected from HudongBaike. The experiment results show that the method is feasible and effective.

Keywords:	knowledge acquisition attribute extraction unstructured text pattern mining
本文献已被 CNKI 万方数据等数据库收录！
	点击此处可从《北京大学学报(自然科学版)》浏览原始摘要信息
	点击此处可从《北京大学学报(自然科学版)》下载免费的PDF全文

设为首页 | 免责声明 | 关于勤云 | 加入收藏