Duke University
5 months to complete at 10 hours a week
This Specialization teaches learners how to create scalable big data pipelines, build machine learning workflows, implement DataOps/DevOps, and develop impactful data visualizations using Python.
Experience in working with Python, Git for version control, Docker for containerization and Kubernetes for deployment and scaling; also a strong foundation in linear algebra and statistics.
The Specialization covers topics such as creating scalable data pipelines using Hadoop, Spark, Snowflake, and Databricks, optimizing data engineering with clustering and scaling, building ML solutions with PySpark and MLFlow, implementing DataOps and DevOps practices, and developing data visualizations with Python.
Duke University is a world-class institution with a strong commitment to applying knowledge in service to society. The instructors for this Specialization are experts in their fields, bringing extensive industry experience and academic expertise to the course content.
This Specialization will equip learners with the skills and knowledge to pursue careers as data-focused software engineers, data scientists, or data engineers working in cloud, machine learning, business intelligence, or other data-driven fields.