Information gathered from internet sources has high variance and different types of noise. This causes out-of-distribution problems with downstream ML modules such as category classification and keyword extraction. The extremely large size of internet-scale datasets requires a solution that is efficient and scalable. The objective of this project is to develop a start-of-the-art anomaly detection […]
Read More