Collaborative dual-PLSA: Mining distinction and commonality across multiple domains for text classification

Fuzhen Zhuang, Ping Luo, Zhiyong Shen, Qing He, Yuhong Xiong, Zhongzhi Shi, Hui Xiong

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

The distribution difference among multiple data domains has been considered for the cross-domain text classification problem. In this study, we show two new observations along this line. First, the data distribution difference may come from the fact that different domains use different key words to express the same concept. Second, the association be-tween this conceptual feature and the document class may be stable across domains. These two issues are actually the distinction and commonality across data domains. Inspired by the above observations, we propose a gen- erative statistical model, named Collaborative Dual-PLSA (CD-PLSA), to simultaneously capture both the domain dis-tinction and commonality among multiple domains. Dif-ferent from Probabilistic Latent Semantic Analysis (PLSA) with only one latent variable, the proposed model has two latent factors y and z, corresponding to word concept and document class respectively. The shared commonality in-tertwines with the distinctions over multiple domains, and is also used as the bridge for knowledge transformation. We exploit an Expectation Maximization (EM) algorithm to learn this model, and also propose its distributed ver-sion to handle the situation where the data domains are geographically separated from each other. Finally, we con-duct extensive experiments over hundreds of classification tasks with multiple source domains and multiple target do-mains to validate the superiority of the proposed CD-PLSA model over existing state-of-the-art methods of supervised and transfer learning. In particular, we show that CD-PLSA is more tolerant of distribution differences.

Original languageEnglish (US)
Title of host publicationHP Laboratories Technical Report
Edition161
StatePublished - 2010

All Science Journal Classification (ASJC) codes

  • Software
  • Hardware and Architecture
  • Computer Networks and Communications

Keywords

  • Clas-sification
  • Cross-domain Learning
  • Statistical Generative Models

Fingerprint

Dive into the research topics of 'Collaborative dual-PLSA: Mining distinction and commonality across multiple domains for text classification'. Together they form a unique fingerprint.

Cite this