基于特征域词频的邮件过滤方法的研究 Research on e-mail filtering by the frequency of the terms in character fields期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

按检索

基于特征域词频的邮件过滤方法的研究

引用本文：	刘慧,马军,雷景生,连莉.基于特征域词频的邮件过滤方法的研究[J].山东大学学报(理学版),2006,41(3):50-53.

作者姓名：	刘慧马军雷景生连莉

作者单位：	1. 山东经济学院,计算机科学与技术学院,山东,济南,250014;山东大学,计算机科学与技术学院,山东,济南,250061 2. 山东大学,计算机科学与技术学院,山东,济南,250061 3. 山东大学,计算机科学与技术学院,山东,济南,250061;海南大学,信息技术学院,海南,海口,570228

基金项目：	山东经济学院校科研和教改项目

摘要：	出了根据邮件特征域信息和特征词频进行垃圾邮件过滤的新方法，并介绍在该方法中的文本特征选取、特征词典构造以及基于TF的权值计算等相关技术，以及改进的文本相似度计算概率模型.实验表明该方法在邮件过滤的查全率、查准率等几个性能评价指标上，比传统的Rocchio方法有了明显改善.
关键词：	垃圾邮件过滤特征域特征词典词频权值计算
文章编号：	1671-9352（2006）03-0134-05
收稿时间：	2006-03-29
修稿时间：	2006年3月29日
Research on e-mail filtering by the frequency of the terms in character fields

LIU Hui,MA Jun,LEI Jing-sheng,LIAN Li.Research on e-mail filtering by the frequency of the terms in character fields[J].Journal of Shandong University,2006,41(3):50-53.

Authors:	LIU Hui MA Jun LEI Jing-sheng LIAN Li

Institution:	School of Computer Science & Technology, Shandong Economic Univ., Jinan 250014, Shandong, China;

Abstract:	A novel method for E-mail filtering is proposed based on the information of character fields and the frequency of the terms in the character fields.The techniques used in the method are discussed,which include selecting the characters of text doc-uments,the constructing the character lexicons as well as the computation of the weights of the term frequency(TF).In addition,an improved probabilistic model for the computation of the similarity of among text documents is provided.Experiments show that the new method is better than traditional Rocchio method in terms of recall,precision and some other evaluation targets.

Keywords:	spare filtering character field character term lexicon term frequency weight calculation
本文献已被 CNKI 维普万方数据等数据库收录！
	点击此处可从《山东大学学报(理学版)》浏览原始摘要信息
	点击此处可从《山东大学学报(理学版)》下载免费的PDF全文

设为首页 | 免责声明 | 关于勤云 | 加入收藏