The Unreasonable Effectiveness of Traditional Information Retrieval in Crash Report Deduplication

Campbell, J.C.; Santos, E.A.; Hindle, Abram

doi:doi:10.7939/r3-jt86-fn23

This decommissioned ERA site remains active temporarily to support our final migration steps to https://ualberta.scholaris.ca, ERA's new home. All new collections and items, including Spring 2025 theses, are at that site. For assistance, please contact erahelp@ualberta.ca.

View

Download

Communities and Collections

Computing Science, Department of / Conference Papers (Computing Science)

Usage

64 views
95 downloads

The Unreasonable Effectiveness of Traditional Information Retrieval in Crash Report Deduplication

Author(s) / Creator(s)
Organizations like Mozilla, Microsoft, and Apple are flooded with thousands of automated crash reports per day. Although crash reports contain valuable information for debugging, there are often too many for developers to examine individually. Therefore, in industry, crash reports are often automatically grouped together in buckets. Ubuntu's repository contains crashes from hundreds of software systems available with Ubuntu. A variety of crash report bucketing methods are evaluated using data collected by Ubuntu's Apport automated crash reporting system. The trade-off between precision and recall of numerous scalable crash deduplication techniques is explored. A set of criteria that a crash deduplication method must meet is presented and several methods that meet these criteria are evaluated on a new dataset. The evaluations presented in this paper show that using off-the-shelf information retrieval techniques, that were not designed to be used with crash reports, outperform other techniques which are specifically designed for the task of crash bucketing at realistic industrial scales. This research indicates that automated crash bucketing still has a lot of room for improvement, especially in terms of identifier tokenization
Date created

2016
Subjects / Keywords
Type of Item

Conference/Workshop Presentation
DOI

https://doi.org/10.7939/r3-jt86-fn23
License

Attribution 4.0 International

Language
- English