What Professionals Need to Know About Data Science and AI, According to Harvard Business School Online

# Introduction
You don't need to be a data scientist to benefit from data science and artificial intelligence (AI). However, you must understand what these technologies can do, where they can fail, and how to evaluate their effects.
After working with data science and AI tools for several years, I've come to realize that people often start with the technology rather than the problem. They want to train a model, launch a chatbot, or build an AI system before deciding what solution they are trying to improve.
For most experts, the goal should not be to know every algorithm. It should be developing enough data and AI knowledge to ask better questions, challenge unreliable results, understand the limitations of these systems, and make informed decisions.
This article includes information from my interview with him Iavor I. BojinovAssociate Professor of Business Administration at Harvard Business School, and my experience working with data science and AI tools. Bojinov co-leads the HBS Online courses on AI for Leaders and Data Science and AI for Decision Making.
# Starting with a decision
Before using AI, I first explain the decision that needs to be made.
“We want to use AI in customer analytics” is too broad. A more specific objective would be: “We want to identify customers who are likely to leave so that our retention team can contact them.”
This links the model to measurable results and actual business action. A highly technologically advanced system is of little value if it does not improve the underlying decision.
Bojinov makes the same point. AI should be considered as a contribution to decision-making rather than a substitute for judgement. Professionals need to understand its capabilities and limitations, validate its recommendations, and focus on better business outcomes rather than embracing AI for its own sake.
In his HBS Online discussion, AI Adoption, Trust, and Decision-Making, Bojinov also explained that successful AI adoption depends on trust and how organizations successfully introduce and measure these systems.
Before starting an AI project, ask what problem it will solve, who will use the output, what action will follow, and how success will be measured.
# Understanding Data First
When I start a data science project, I don't immediately train a model. I start by examining the data.
Exploratory data analysis (EDA) helps us understand dataset structure, quality, distribution, and relationships. It may reveal missing values, duplicate records, inconsistent categories, unusual observations, and potential bias.
These checks are important because accuracy alone can be misleading.
Let's say only five percent of customers leave. A model that predicts each customer will always be 95 percent accurate, yet will not identify a single customer at risk.
Professionals should therefore question whether the metric reflects the true business objective rather than relying on a single scheme that looks good.
The latest HBS Online article, Data Preparation for Your AI Model: Importance and Best Practicesand emphasizes that the performance of a model is highly dependent on the quality of the data behind it. The article covers cleaning, enriching, transforming, and organizing data prior to model development.
# Careful Data Preparation
Real-world data is rarely suitable for machine learning.
Data processing may involve handling missing values, correcting formats, removing duplicates, transforming categories, selecting useful features, and separating training data from test data.
Professionals must also know where their data came from, how it was collected, what may be missing, and whether previous decisions introduced bias.
In my experience, improving the data often brings more benefits than replacing a simple algorithm with a more complex model.
The model learns from the information it receives. If that information is incomplete, out of date, inconsistent, or biased, the results will also be unreliable.
Better technology cannot fully compensate for bad data.
# Choosing the Simplest Model That Works
The best model isn't always the newest, biggest, or most expensive.
For structured business data — such as customer records, transactions, sales history, or performance data — a regression model or decision tree may be more appropriate than a general-purpose language model (LLM).
It can also be faster, cheaper, easier to define, and easier to monitor.
LLMs are useful if the work involves language, such as summarizing documents, extracting information, producing drafts, or analyzing customer feedback. For forecasting, planning, and tabular data, traditional machine learning may always be a better option.
Before choosing a model, I compare predictive performance, training costs and indicators, speed, scaling, interpretation, maintenance requirements, and the cost of incorrect predictions.
A more accurate model is not always the best choice if it is more expensive, difficult to define, or difficult to maintain.
The goal is not to use the most advanced AI. It's about choosing a reliable and cost-effective solution to a problem.
# Verification Before Trust
A model that works well during development does not guarantee that it will work in the real world.
It should be tested using data that was not used to train or select it. Test data should also be representative of the customers, markets, and situations in which the model will operate.
Verification should continue after submission. Changes in customer behavior, market shifts, disruptions in data pipelines, and patterns that were once useful may become less reliable.
Bojinov points to weak validation and testing as a common error. If organizations use models without properly testing them, they may be overconfident about the results.
Experts should ask how the model was tested, what errors it makes, whether its test data is representative, and how its performance will be monitored over time.
Trust should come from evidence, not from how advanced or confident a system appears to be.
# Converting Customer Insights into Action
Customer statistics should be above the dashboard.
The dashboard can show what happened. A predictive model can estimate what happens next. An assessment can help determine what the organization should do about it.
Predicting customer churn is only important if the company has a realistic retention strategy. It should then measure whether the intervention actually works.
Experts should also avoid confusing correlation with causation. The factor associated with customer churn is not the reason customers leave.
This is where exploratory thinking becomes valuable. Rather than assuming that an intervention will work, organizations should test it, measure the effect, and learn from the evidence.
The goal is not to generate more predictions. It's about making better decisions.
# Managing LLMs as assistants
Major linguistics majors (LLMs) have made AI more accessible. I use them regularly for coding, researching, improving grammar, summarizing information, generating ideas, and learning something new.
They help me work quickly and explore unfamiliar topics. However, fluent and confident language is not the same as factual accuracy.
LLMs misunderstand the context, rely on out-of-date information, create sources, generate incorrect code, or ignore security risks. HBS Online Guide, What Are the Major Language Patterns and How Do They Work?provides an overview of what business professionals should understand before using them.
I never trust their result blindly. I review and re-read their outputs, verify key claims, test the generated code, check for functionality, and check for security issues.
I treat the LLM as an assistant rather than an unquestioned authority. The greater the results of the work, the more carefully its result should be reviewed.
# Keeping Human Judgment in the Loop
According to Bojinov, knowledge of data literacy, analytical thinking, and the ability to critically evaluate the effects of AI will become more important as these tools become more widely used.
Experts don't need to be AI experts, but they do need to be fluent enough to ask the right questions, interpret the evidence, spot weak or unreliable results, and understand where the answer is unreliable.
HBS Online The Parlor Room episode, How Mid-Time Professionals Can Earn and Grow with AIexplores how professionals can combine AI fluency with their existing domain knowledge to create value in an AI-driven workplace. Another episode, Early Career Advice for Building AI and Human Skillsdiscusses the importance of developing strong technology while learning to work effectively alongside AI.
Ultimately, the professionals who will stand out will be those who can combine technical understanding with domain knowledge and sound business judgment.
# Final thoughts
Many companies are rushing to add AI everywhere because they don't want to be left behind. By doing so, they may spend a lot of money on APIs, infrastructure, training, and maintenance without questioning whether AI is really the right solution.
I think the hype will end up being right.
Companies will learn when AI adds real value, when people always matter, and when a simple and cheap solution is enough.
Successful data science and AI projects don't start with the latest model. They start with a clear problem, reliable data, proper evaluation, realistic costs, and a plan to turn results into action.
The goal is not to use AI everywhere. It's about using it deliberately, validating it carefully, and using it where it really improves decisions.
Abid Ali Awan (@1abidiawan) is a data science expert with a passion for building machine learning models. Currently, he specializes in content creation and technical blogging on machine learning and data science technologies. Abid holds a Master's degree in technology management and a bachelor's degree in telecommunication engineering. His idea is to create an AI product using a graph neural network for students with mental illness.



