Development of a Big Data Processing and Analytics Pipeline
Data Science — 2025, Undergraduate
This study developed a big data processing and analytics pipeline that ingests large files, splits workloads, and produces summary analytics with progress tracking. The pipeline validates schemas and schedules batch runs. Adopting a client-server architecture and a survey-based usability evaluation, the system was built using Python, Django, and PostgreSQL. The findings showed that the pipeline handled large datasets more efficiently than manual methods, while achieving a high System Usability Scale score. The study recommends deployment in data-intensive research units.
Homepage screenshot
