Guide · 6 min read

Regression Analysis for Business Assignments

Regression shows how an outcome moves with one or more drivers. Learn to calculate a simple regression by hand, read the output table and describe the result in business terms.

What regression does

Regression fits a line (or a plane, with more variables) through data to describe and predict a numerical outcome. The outcome is the dependent variable, usually called y. The drivers are independent variables, called x. A simple regression has one x, such as advertising spend. A multiple regression has several, such as advertising spend, price and season.

The simple model is y = b0 + b1 x. The intercept b0 is the predicted y when x is zero. The slope b1 is the predicted change in y for a one-unit increase in x. Least squares chooses the line that makes the sum of squared vertical gaps between the data and the line as small as possible.

Worked example: calculating a simple regression by hand

A business records monthly advertising spend and sales, both in thousands of dollars, for six months.

MonthAdvertising xSales y
1220
2324
3429
4531
5638
6740

Step 1, means. Mean of x is 4.5. Mean of y is 182 / 6 = 30.33.

Step 2, sums of squares and products. The sum of (x minus mean x) squared is 6.25 + 2.25 + 0.25 + 0.25 + 2.25 + 6.25 = 17.5. The sum of (x minus mean x) times (y minus mean y) is 25.83 + 9.50 + 0.67 + 0.33 + 11.50 + 24.17 = 72.0.

Step 3, slope and intercept. b1 = 72.0 / 17.5 = 4.114. b0 = 30.33 - 4.114 x 4.5 = 11.82.

The fitted line is sales = 11.82 + 4.114 x advertising.

xActual yPredicted yResidual (actual minus predicted)
22020.05-0.05
32424.16-0.16
42928.270.73
53132.39-1.39
63836.501.50
74040.62-0.62

The residuals add to approximately zero, as they should. The sum of squared residuals (SSE) is about 5.13. The total variation in sales around its mean (SST) is 301.33. So R-squared = 1 - 5.13 / 301.33 = 0.983. The standard error of the estimate is the square root of 5.13 / 4, which is 1.13.

Interpreting the result in business terms

  • Slope: each additional $1,000 of advertising is associated with about $4,114 more in sales, on average, in the range of data observed.
  • Intercept: the model predicts sales of about $11,820 with no advertising. Be careful, because zero advertising is outside the observed range of 2 to 7, so this is an extrapolation and not a reliable claim.
  • R-squared: advertising explains about 98 percent of the variation in sales across these six months. That is very high and unusual in real data. With only six points, treat it cautiously.
  • Prediction: with advertising of 5.5, predicted sales are 11.82 + 4.114 x 5.5 = $34.45 thousand, which is inside the data range, so this is a more trustworthy interpolation.

Say associated with rather than caused by unless the data come from an experiment. Regression describes patterns, and it does not by itself prove that advertising causes sales. Seasonality or pricing could drive both.

Reading Excel output

Excel's regression tool (Data Analysis ToolPak) produces three blocks. Here is the output for the example, with explanations.

Regression statisticsValueMeaning
Multiple R0.991Correlation between actual and predicted y
R Square0.983Share of variation in y explained
Adjusted R Square0.979R-squared adjusted for the number of predictors, better for comparing models
Standard Error1.13Typical size of prediction errors, in y units
Observations6Sample size
CoefficientsCoefficientStandard errort StatP-value
Intercept11.821.309.070.0008
Advertising4.1140.27115.2less than 0.001

The t statistic is the coefficient divided by its standard error (4.114 / 0.271 = 15.2). It tests whether the true slope is zero. A small p-value, here far below 0.05, means advertising has a statistically significant relationship with sales. The F-test in the ANOVA block tests whether the model as a whole is useful, and with one predictor it is the square of the slope's t statistic (15.2 squared is about 231). The 95 percent confidence interval for the slope is roughly 4.114 plus or minus 2.78 x 0.271, which is from about 3.36 to 4.87.

Working on this assignment now? Get a price for help with your paper.

Get an instant quote

Multiple regression

Adding variables lets you separate effects. Suppose sales depend on advertising and price. The model is sales = b0 + b1 x advertising + b2 x price. Interpretation changes slightly: b1 is the change in sales for a one-unit increase in advertising holding price constant, and b2 is the effect of price holding advertising constant.

FeatureHow to handle it
Categorical predictors (for example holiday month yes or no)Create a dummy variable coded 0 or 1; its coefficient is the difference in y between the two groups, holding the other variables constant
More than two categoriesUse one fewer dummy than the number of categories; the left-out category is the baseline
Adjusted R-squaredUse it rather than R-squared to compare models with different numbers of variables, since R-squared never falls when you add variables
MulticollinearityWhen predictors are highly correlated with each other, coefficients become unstable; check correlations or variance inflation factors (VIF above about 5 to 10 is a warning)
Which variables to keepKeep those with sensible reasoning and statistical significance, not just a high R-squared

Assumptions and how to check them

AssumptionWhat it meansHow to checkIf violated
LinearityThe relationship between x and y is a straight lineScatter plot, residuals versus fitted valuesTransform variables or add a curved term
IndependenceErrors are not related to each otherPlot residuals over time; Durbin-Watson test for time seriesUse time-series methods
Constant varianceThe spread of residuals is similar across fitted valuesResidual plot with a fan shape is a warningTransform y or use robust methods
Normal errorsResiduals are roughly normalHistogram or normal probability plotMatters mostly for small samples

In an assignment, include a scatter plot and a residual plot, and write one or two sentences on whether the assumptions look reasonable. Also check for outliers and influential points, because a single unusual observation can change the line a lot.

How to write up regression results

A short write-up (using the example)

A simple linear regression of monthly sales on advertising spend, using six months of data, found a significant positive relationship (b = 4.11, t(4) = 15.2, p < .001). Each additional $1,000 spent on advertising was associated with about $4,100 in extra sales. The model explained 98 percent of the variance in sales (R-squared = .98). Because the sample is small and covers a limited range of spending (from $2,000 to $7,000), the results should not be extrapolated beyond that range, and the data do not show that advertising causes the increase in sales.

  • State the model Say what y and x are, and the number of observations.
  • Report the key statistics Coefficients, significance, R-squared and sample size.
  • Interpret in business terms Use units and plain language, and say what a change in x means for y.
  • Check and mention assumptions Include at least one diagnostic plot.
  • Avoid overclaiming Use associated with unless causation can be justified, and avoid extrapolation.

Common mistakes

  • Confusing correlation with causation A strong relationship does not show that x causes y.
  • Extrapolating Predicting far outside the range of the data is unreliable.
  • Relying on R-squared alone A high R-squared does not mean the model is correct. Check residuals.
  • Ignoring units Always interpret the slope in the units of x and y.
  • Adding variables without thought Each variable should have a reason, not just improve the fit.
  • Forgetting categorical variables need dummies Do not enter categories as numbers 1, 2, 3.

If you want help running regressions or writing up results, you can order a data analysis report and upload your data.

Quick answers

What is a good R-squared?

It depends on the field. In controlled processes it can be very high. In human behavior, 0.2 to 0.5 can be useful. Judge the model by its purpose and its diagnostics, not a fixed number.

What does the p-value of a coefficient tell me?

It tests whether the true coefficient is zero. A small p-value suggests the variable has a real relationship with y after accounting for the other variables in the model.

When should I use adjusted R-squared?

When comparing models with different numbers of predictors, because adjusted R-squared penalizes variables that add little explanatory power.

Can I use regression with a yes or no outcome?

A standard linear regression is not appropriate. Use logistic regression for binary outcomes.

Need a hand with your paper?

Tell us the assignment and see your price straight away.

Get an instant quote