基于树形结构的Web信息抽取 Web Information Extraction Based on Tree Structure期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

按检索

基于树形结构的Web信息抽取

引用本文：	任仲晟,薛永生.基于树形结构的Web信息抽取[J].福建师范大学学报(自然科学版),2009,25(3).

作者姓名：	任仲晟薛永生

作者单位：	1. 福建师范大学数学与计算机科学学院,福建,福州,350108 2. 厦门大学计算机科学系,福建,厦门,361005

基金项目：	国家自然科学基金，福建省自然科学基金，福建省重点科技项目

摘要：	提出了一种基于树形结构的Web结构化数据抽取算法.该算法基于HTML的树形层次结构,包括HTML树构造算法,数据区域挖掘算法,数据记录挖掘算法以及数据记录模式生成算法.算法引入了页面元素布局位置等信息用于清洗页面,采用层次划分思想实现页面数据区域的挖掘,并通过树匹配生成记录模式,实现最终数据项抽取.实验表明,该方法可以有效地实现Web结构化数据抽取.
关键词：	Web数据抽取 Web挖掘信息抽取
Web Information Extraction Based on Tree Structure

REN Zhong-sheng,XUE Yong-sheng.Web Information Extraction Based on Tree Structure[J].Journal of Fujian Teachers University(Natural Science),2009,25(3).

Authors:	REN Zhong-sheng XUE Yong-sheng

Institution:	1.School of Mathematics and Computer Science;Fujian Normal University;Fuzhou 350108;China;2.Department of Computer Science;Xiamen University;Xiamen 361005;China

Abstract:	It proposes tree structure based Web data extraction algorithm in view of the inadequacies of the existing methods.The tree structure based algorithm includes: the algorithm of HTML tree construction,the algorithm of data region mining,the algorithm of data record mining,and the algorithm of record schema generation.The algorithm cleans the Web pages using the position information of page elements,mines data region by hierarchical clustering, and generates record schema finishing data item extraction throug...

Keywords:	Web data extraction Web mining information extraction
本文献已被 CNKI 维普万方数据等数据库收录！

设为首页 | 免责声明 | 关于勤云 | 加入收藏