一种基于标签的改进主题演化模型

设为首页

收藏本站

网站地图 | English | 公务邮箱

读者指南

学术客户端

NSTL服务站

科技查新

一种基于标签的改进主题演化模型

详细信息查看全文 | 推荐本文 |

英文篇名：An Improved Topics over Time Model Based on Label
作者：姚立 ; 张曦煌
英文作者：YAO Li;ZHANG Xihuang;School of Internet of Things Engineering,Jiangnan University;
关键词：标签 ; 主题演化模型 ; 隐狄利克雷分配 ; 词频-反重力距算法 ; 吉布斯采样
英文关键词：label;;Topics over Time(ToT) model;;Latent Dirichlet Allocation(LDA);;TF-IGM algorithm;;Gibbs sampling
中文刊名：JSJC
英文刊名：Computer Engineering
机构：江南大学物联网工程学院;
出版日期：2018-04-04 10:35
出版单位：计算机工程
年：2019
期：v.45;No.499
基金：江苏省产学研合作项目(BY2015019-30)
语种：中文;
页：JSJC201904034
页数：7
CN：04
ISSN：31-1289/TP
分类号：211-216+222

摘要

传统主题演化(ToT)模型通常忽略原始数据中的标签元信息。为此,建立一种基于标签的改进ToT模型。针对传统权重算法忽略词汇在文档集类别间和类别内的分布对权重产生影响的问题,结合文档标题特征,使用改进词频-反重力距算法进行权重分析,以扩展模型的生成过程。在ToT模型的基础上引入原始文档的标签属性,构建改进模型并使用吉布斯采样算法估计其参数。实验结果表明,与ToT模型相比,该模型具有较高的泛化能力。
Traditional Topics over Time(ToT) models usually ignore label meta-information in the original data.To solve this problem,an improved ToT model based on label is established.Aiming at the problem that traditional weighting algorithms ignore the influence of vocabulary distribution among and within document sets on weights,combined with the characteristics of document titles,an improved TF-IGM algorithm is used to analyze the weights to extend the generation process of the model.Based on the ToT model,the label attributes of the original document are introduced to construct the improved model and estimate its parameters using Gibbs sampling algorithm.Experimental results show that the proposed model has higher generalization ability than the ToT model.

引文

[1] BLEI D M,NG A Y,JORDAN M I.Latent Dirichlet allocation[J].Journal of Machine Learning Research,2003,3:993-1022.
    [2] BLEI D M.Probabilistic topic models[J].Communications of the ACM,2012,55(4):77-84.
    [3] BLEI D M,LAFFERTY J D.Dynamic topic models[C]//Proceedings of the 23rd International Conference on Machine Learning.New York,USA:ACM Press,2006:113-120.
    [4] WANG X,MCCALLUM A.Topics over time:a non-Markov continuous-time model of topical trends[C]//Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.New York,USA:ACM Press,2006:424-433.
    [5] 刘良选,黄梦醒.一种面向词汇突发的连续时间主题模型[J].计算机工程,2016,42(11):195-201.
    [6] RAMAGE D,HALL D,NALLAPATI R,et al.Labeled LDA:a supervised topic model for credit attribution in multi-labeled corpora[C]//Proceedings of 2009 Con-ference on Empirical Methods in Natural Language Processing.New York,USA:ACM Press,2009:248-256.
    [7] RAMAGE D,MANNING C D,DUMAIS S.Partially labeled topic models for interpretable text mining[C]//Proceedings of ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.New York,USA:ACM Press,2011:457-465.
    [8] RAMAGE D,HEYMANN P,MANNING C D,et al.Clustering the tagged Web[C]//Proceedings of the 2nd ACM International Conference on Web Search and Data Mining.New York,USA:ACM Press,2009:54-63.
    [9] 张小平,周雪忠,黄厚宽,等.一种改进的LDA主题模型[J].北京交通大学学报,2010,34(2):111-114.
    [10] SOUCY P,MINEAU G W.Beyond TFIDF weighting for text categorization in the vector space model[C]//Proceedings of the 19th International Joint Conference on Artificial Intelligence.San Francisco,USA:Morgan Kaufmann Publishers Inc.,2005:1130-1135.
    [11] CHEN K,ZHANG Z,LONG J,et al.Turning from TF-IDF to TF-IGM for term weighting in text classification[J].Expert Systems with Applications,2016,66:245-260.
    [12] 侯汉清,章成志,郑红.Web概念挖掘中标引源加权方案初探[J].情报学报,2005,24(1):87-92.
    [13] 张宏毅,王立威,陈瑜希.概率图模型研究进展综述[J].软件学报,2013,24(11):2476-2497.
    [14] BOWMAN K O,SHENTON L R.Parameter estimation for the Beta distribution[J].Journal of Statistical Computation and Simulation,2007,43(3/4):217-228.
    [15] GUO X,XIANG Y,CHEN Q,et al.LDA-based online topic detection using tensor factorization[J].Journal of Information Science,2013,39(4):459-469.
    [16] WEI X,CROFT W B.LDA-based document models for ad-hoc retrieval[C]//Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval.New York,USA:ACM Press,2006:178-185.

常见问题　|　交通位置　|　联系我们　|　OA远程办公

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700