Academic Commons

Presentations (Communicative Events)

A description of the CIDR system as used for TDT-2

Radev, Dragomir R.; McKeown, Kathleen; Hatzivassiloglou, Vasileios

We describe several experimental parameters and a parallelization technique used in our online document clustering system, CIDR. These modifications were introduced into CIDR to reduce the running time so that incoming documents be clustered in almost real time. We discuss how several of these parameters are justified on linguistic grounds and report preliminary quantitative results on the effects that these parameters have on speed and accuracy.

Files

More About This Work

Academic Units
Computer Science
Publisher
DARPA Broadcast News Workshop
Published Here
May 3, 2013
Academic Commons provides global access to research and scholarship produced at Columbia University, Barnard College, Teachers College, Union Theological Seminary and Jewish Theological Seminary. Academic Commons is managed by the Columbia University Libraries.