EECS 280, Programming and Introductory Data Structures
Two-Sample Statistics Tool
A command-line tool that splits a CSV column into two groups, prints descriptive statistics for each, and estimates a 95% confidence interval for the difference in means using bootstrap resampling.
What I built
- Implemented a statistics library over std::vector: count, sum, mean, median, min, max, corrected sample standard deviation, and an Excel-style interpolated percentile.
- Built a two-sample analysis driver that filters a CSV data column by a second column's value and reports stats for groups A and B.
- Approximated the sampling distribution of the mean difference with 1000 bootstrap resamples and derived the 95% interval from the 2.5th and 97.5th percentiles.
- Wrote unit tests for the library and checked the driver against a reference output with a Makefile regression target built under AddressSanitizer and UBSan.