Agile based processes have proven to deliver valuable and high quality traditional software based solutions to business users in a timely manner. Over the last few years, data science related projects are becoming more and more important in every industry. Software engineering has tried different development methodologies - waterfall, Scrum, XP, etc. and data science can learn from the experiences of the software engineering.
In this post, I will try to identify the most important and applicable values, principles and practices from agile that can be used in a typical data science project. We will start with a brief overview of software engineering, agile, data science and conclude with some recommendations on how to incorporate agile into data science projects.
Because of the sequential nature of the waterfall approach, the individual roles of the team members are fixed.
Agile teams are cross functional in nature having all skills required to deliver and deploy a working software product. Also, there is a flexibility in individual roles with team members playing multiple roles as needed by the team.
Though the Cross Industry Standard Process for Data Mining (CRISP-DM)[5] was developed 20 years ago, its concepts and phases are still relevant for data science projects. Microsoft is defining a team process for data science that's based on CRISP-DM[7]. Though CRISP-DM does talk about these phases are cyclical in nature but some projects may end up following them in a sequential and waterfall approach.
Roger Peng talks about iterative process that is applied to all steps of the data analysis in his "Epicycles of Analysis" [8].
Similarly, there have been attempts to define values and principles for data science [3, 6] similar to agile values and principles [1].
Though software engineering and data science are two different disciplines with each requiring different set of skills, one could draw parallels between software engineering and data science. Most of the data science project deliverables can be considered as data products offering insights for the business stakeholder. Data science projects are typically non-linear in nature, one may learn something along the way that requires changes. Agile methodologies have built in practices and process to allow for such changes.
In this post, I will try to identify the most important and applicable values, principles and practices from agile that can be used in a typical data science project. We will start with a brief overview of software engineering, agile, data science and conclude with some recommendations on how to incorporate agile into data science projects.
Software Engineering
Before agile or iterative software development methodologies, software engineering projects typically followed waterfall approach where activities are scheduled one after the other. If during testing or after deployment if the product was found not to meet the expectations of the business users, the whole process is repeated in sequence.![]() |
| Waterfall Approach |
Because of the sequential nature of the waterfall approach, the individual roles of the team members are fixed.
Agile
In agile development methodologies[1,2], teams go through the traditional software engineering activities in a time boxed (typically 2 weeks) iteration and a software project will have multiple such iterations. What's different from the waterfall approach is that at the end of every iteration, the team delivers/deploys something of value to the business. This allows for the business and team to make adjustments along the way while delivering something of value every iteration.![]() |
| Iterative Agile Approach |
Agile teams are cross functional in nature having all skills required to deliver and deploy a working software product. Also, there is a flexibility in individual roles with team members playing multiple roles as needed by the team.
Data Science
A typical data science project, goes through following stages. Just like in traditional software engineering projects, there is a tendency to do these activities in sequence: acquire/prepare/cleanse all data needed before exploring data or building models.![]() |
| Typical Data Science Pipeline |
Though the Cross Industry Standard Process for Data Mining (CRISP-DM)[5] was developed 20 years ago, its concepts and phases are still relevant for data science projects. Microsoft is defining a team process for data science that's based on CRISP-DM[7]. Though CRISP-DM does talk about these phases are cyclical in nature but some projects may end up following them in a sequential and waterfall approach.
![]() |
| CRISP-DM Process Diagram by Kenneth Jensen |
![]() |
| From The Art of Data Science, Roger D. Peng and Elizabeth Matsui |
Similarly, there have been attempts to define values and principles for data science [3, 6] similar to agile values and principles [1].
Agile Data Science
"In preparing for battle, I have always found that plans are useless, but planning is indispensable."
–General Dwight D. Eisenhower
Though software engineering and data science are two different disciplines with each requiring different set of skills, one could draw parallels between software engineering and data science. Most of the data science project deliverables can be considered as data products offering insights for the business stakeholder. Data science projects are typically non-linear in nature, one may learn something along the way that requires changes. Agile methodologies have built in practices and process to allow for such changes.
Values
Values are fundamental and are guidelines for the behavior of team members. Agile manifesto [1] values:
- Individuals and interactions over processes and tools
- Working software over comprehensive documentation
- Customer collaboration over contract negotiation
- Responding to change over following a plan
Data Science manifesto [2] values:
- Minimal Viable Products over prototypes
- APIs over databases
- Clever use of computation over convenient assumptions
- Dashboards over reports
- Validation, scrutiny and repeatability over convention and ad verecundiam
Important values to focus on:
- The concept of Minimal Viable Data Product (MVDP similar to MVP in Software Engineering) that is working and delivering insight/value to business is something that's central and important.
- Reproducibility which allows independent validation is critical to data science projects.
- Working model
- Collaborate and interact with end user frequently
- Plan ahead but accept change
- Automate as much as possible
![]() |
| Minimal Viable Product (MVP) |
Principles
Principles can be described as rules or "truths", arising from experience, knowledge, and (often) values. Principles are guides to behavior.
Agile principles[2]:
- Satisfy the customer through early and continuous delivery.
- Welcome changing requirements
- Deliver working software frequently
- Business people and developers must work together
- Build projects around motivated individuals
- Team face-to-face conversation.
- Working software is the primary measure of progress.
- Agile processes promote sustainable development
- Technical excellence and good design enhances agility.
- Simplicity--the art of maximizing the amount of work not done--is essential.
- The best architectures, requirements, and designs emerge from self-organizing teams.
- Reflect at regular intervals.
Data Science principles[2]:
- Aim to completely remove manual intervention (automate).
- Data science is about solving problems, not models or algorithms.
- All validation of data, hypotheses and performance should be tracked, reviewed and automated.
- Prior to building a model, construct an evaluation framework with end-to-end business focused acceptance criteria.
- A product needs a pool of measures to evaluate its quality.
- Even research can be broken down into clearly defined tasks.
- Don’t neglect assumptions in models.
Some important principles to adopt:
- Reflect at regular intervals and adjust
- Automate as much as possible
- Focus and solve business problems
- Deliver and demonstrate working models frequently
- Simplicity in design and models
- Accept change and learn
Practices
Agile methodologies strongly recommend good team and engineering practices that allow principles and values being acted and followed.
![]() |
| Product, Sprint (Iteration) Backlog |
Agile practices:
- Time boxed iteration (or sprint of 2/4 weeks) to deliver working model that provides value
- Backlog of priority items to work on next
- Review backlog frequently to prioritize
- Iteration planning
- Daily stand up - quick 10 minute review of progress and any blocks
- Team board to show status
- Source control and continuous check ins to allow other team members to use one's work
- Automated builds
- Continuous deployment
- Work in pairs (four set of eyes)
- Iteration retrospective to review what worked and what didn't work
Data science project will definitely benefit by implementing some of the agile practices:
- Source control with hosted Notebooks
- Backlog of priority items
- Regular review of iteration (daily)
- End of iteration retrospective
- Status board to show team's progress
- Automated processes for builds, deployments
- Snapshots/versions of data (training data, test data)
Tools
- Source control (Git, SVN)
- Agile process tools like JIRA
- Notebooks (e.g., R Notebook, Jupyter Python Notebook, Zeppelin) are good for self-documentation, reproducibility with code and also with visualizations that both data scientists and business users have access to. Hosted notebooks are preferred.
- Good data pipeline engineering platforms/workbenches for big data platforms (e.g., CASK, Cloudera, Trifacta)
References
- Agile Manifesto: Values and Principles
- Agile Principles and Values, by Jeff Sutherland, MSDN
- Data Science Manifesto
- Successful Data Teams are Agile and Cross-Functional, John Akerd
- CRISP-DM
- Ten Simple Rules for Effective Statistical Practice, Kass RE, Caffo BS, Davidian M, Meng X-L, Yu B, Reid N.
- Team Data Science Process, Microsoft
- The Art of Data Science, Roger D. Peng, Elizabeth Matsui







What a fantastic read on Data Science. This has helped me understand a lot in Data Science course. Please keep sharing similar write ups on Data Science. Guys if you are keen to know more on Data Science, must check this wonderful Data Science tutorial and i'm sure you will enjoy learning on Data Science training.:-https://www.youtube.com/watch?v=IJ9sdk_Sc3g
ReplyDeleteKadangpintar | Online Casino and Sport Betting
ReplyDeleteKadangpintar, online casino and sports betting is a popular destination. Discover everything you need to know about Kadangpintar in our 인카지노 latest 제왕카지노 Online Casino 온카지노