1999 Presentations (Communicative Events)
A description of the CIDR system as used for TDT-2
We describe several experimental parameters and a parallelization technique used in our online document clustering system, CIDR. These modifications were introduced into CIDR to reduce the running time so that incoming documents be clustered in almost real time. We discuss how several of these parameters are justified on linguistic grounds and report preliminary quantitative results on the effects that these parameters have on speed and accuracy.
Subjects
Files
- radev_al_99a.pdf application/pdf 70.1 KB Download File
More About This Work
- Academic Units
- Computer Science
- Publisher
- DARPA Broadcast News Workshop
- Published Here
- May 3, 2013