From OLS to Negative Binomial: Helping a UNT Political Science PhD Candidate Fix Count Data Analysis in SPSS
SPSS negative binomial regression consulting is what a Political Science PhD candidate at the University of North Texas needed after her committee rejected her entire results chapter. Her theory was sound. Her literature review was clean. But her model was wrong, and no one had told her why.
Note: names and some identifying details in this case have been altered to protect the student’s privacy, but the data problem and the fix are unchanged.
I have over 20 years of analytics consulting experience across R, Python, SPSS, Stata, and Power BI, and I hold an MBA in Finance and an M.Tech. I mention this not to pad my bio, but because this case is a good example of something I keep seeing in social science dissertations. Across countries and universities, regardless of how strong the underlying theory is, the statistics often get bolted on at the end. Nobody checks whether the model actually fits the data type.
What Is Count Data, And Why OLS Regression Fails It
Count data analysis in SPSS starts with recognising what kind of variable you actually have. Count data is any outcome that counts the number of times something happens. Number of protest events per city. Number of hospital visits per patient. Number of customer complaints per month. These numbers cannot go below zero, and they are often bunched near zero with a long tail of larger values.
Negative binomial regression is a statistical model used for count data that shows overdispersion, meaning the variance is larger than the mean. It extends Poisson regression by adding a dispersion parameter that accounts for this extra variability, an approach formalised in Cameron and Trivedi’s Regression Analysis of Count Data (Cambridge University Press), one of the standard references on this topic. IBM SPSS Statistics implements this model through its Generalized Linear Models procedure, which is the tool we used in this case.
OLS regression assumes your outcome variable is continuous and roughly normally distributed. Count data almost never behaves this way.
Why does count regression give negative predicted values?
This happens when OLS is applied to a variable that can only be zero or positive. OLS has no built-in floor at zero, so with skewed count data and many low values, the fitted line can dip below zero for some observations. That is a mathematical impossibility for a count outcome, and it is usually the first clue that OLS is the wrong model, not that something is wrong with the data itself.
Signs Your Dependent Variable Is Count Data
- The variable only takes whole numbers (0, 1, 2, 3…)
- A large share of observations are zero
- The distribution is skewed to the right, not bell shaped
- You are counting events, occurrences, or incidents over a fixed period
If two or more of these apply to your dissertation variable, OLS is probably the wrong choice. You can check your regression assumptions before your committee does it for you, using the regression assumption checker.
In my opinion, this is the one check every methods course should teach in week one, and almost none of them do. Students are taught to run OLS by default because it is the model everyone learns first, not because anyone stops to ask whether it fits the outcome variable.
Understanding Overdispersion In SPSS
Overdispersion in SPSS happens when the variance of your count variable is much bigger than its mean. Poisson regression assumes these two are roughly equal. When they are not, Poisson underestimates your standard errors, and suddenly variables look statistically significant when they are not.
You can check this using the deviance-to-degrees-of-freedom ratio from your Poisson output. A ratio close to 1 means Poisson is fine. A ratio well above 1, like the value over 4 I found in this student’s data, confirms severe overdispersion.
Poisson vs Negative Binomial vs Zero-Inflated Models
| Feature | Poisson | Negative Binomial | Zero-Inflated |
|---|---|---|---|
| Assumes variance = mean | Yes | No | No |
| Handles overdispersion | No | Yes | Depends on base model |
| Handles excess zeros | No | Partially | Yes, directly |
| Best for | Clean, low-variance counts | Real-world messy counts with many zeros | Two distinct processes generating zeros (e.g. “never protests” cities vs “protests rarely”) |
Zero-inflated models are worth a separate mention because students often confuse them with negative binomial regression. Negative binomial handles extra variance across the whole distribution. Zero-inflated models are for a specific situation: when zeros come from two different sources, for example some cities that structurally never have protests versus cities that simply had a quiet month.
If your zeros have one plausible explanation, negative binomial is usually enough. If they have two competing explanations, a zero-inflated model deserves a look.
If you are unsure which family fits your data, this is exactly the kind of decision I walk students through as part of my SPSS data analysis tutoring work, and it is one of the most common reasons political science PhD statistical consultant queries land in my inbox.
Real Case: A UNT Political Science PhD Candidate’s Model Problem
Her dependent variable was the number of protest events per city-month. Her independent variables included local media coverage and police response. It was a genuinely interesting research question, and this is the part of UNT dissertation SPSS help requests that frustrates me most: good research questions get undermined by a mechanical, fixable error.
Her committee flagged two problems. One member asked why the model produced negative predicted values for a count variable. Another said the results looked unstable across specifications.
When she sent me her codebook, syntax, and output, the issue was clear within minutes. She had run OLS because it was the model she knew best, and she had never included an offset for city population, so larger cities were showing more events purely because they had more people, not because her theory was working.
Step-By-Step: Running Negative Binomial Regression In SPSS
This is the exact process I walked her through, using SPSS’s Generalized Linear Models procedure. I prefer GENLIN over the older NBREG command because it gives more direct control over the offset and the standard-error correction, and its output maps more cleanly onto how reviewers expect GLM results to be reported.
- Go to Analyze > Generalized Linear Models > Generalized Linear Models.
- In the Type of Model tab, select Negative binomial with log link.
- In the Response tab, set your count variable as the dependent variable.
- In the Predictors tab, add your independent variables, for example media coverage, police response, and city demographics.
- In the Offset tab, add the natural log of your exposure variable, for example log of city population, so the model accounts for differences in exposure.
- In the Statistics tab, request parameter estimates, exponentiated estimates, and a corrected covariance estimator to guard against mild remaining misspecification.
- Compare AIC and BIC between the Poisson and negative binomial versions to confirm which fits better.
Here is the SPSS syntax we used for her model:
GENLIN protest_count WITH media_coverage police_response
/MODEL media_coverage police_response INTERCEPT=YES
DISTRIBUTION=NEGBIN LINK=LOG
/PRINT SOLUTION (EXPONENTIATED)
/CRITERIA METHOD=FISHER SCALE=PEARSON COVARIANCE=ROBUST
/OFFSET=log_pop
/MISSING CLASSMISSING=EXCLUDE.
Note: ROBUST here is SPSS’s fixed command keyword for the corrected covariance estimator. It is a required syntax term, not a stylistic word choice.
If your model still runs on OLS logic at this stage, compare it side by side against a proper linear specification using the multiple linear regression calculator to see exactly where the assumptions break down, or use the Poisson distribution calculator to sanity check your baseline count distribution first.
Incidence Rate Ratio Interpretation: What An IRR Actually Means
Once you exponentiate the coefficients from a negative binomial model, you get incidence rate ratios, not raw effect sizes. This is where I see most students get stuck in their write-up, because IRRs read differently from regular regression coefficients.
For her media coverage variable, the IRR came out to 1.34. This means a one-unit increase in media coverage was associated with a 34 percent increase in the expected number of protest events, holding other variables constant. An IRR below 1 means a decrease, and an IRR of exactly 1 means no effect at all.
Comparing Model Fit: AIC, BIC, And Likelihood Ratio Tests
We compared the negative binomial model against the Poisson model using AIC and BIC, and the negative binomial version fit noticeably better on both measures. Lower values on both indicate better fit, so this gave her objective, defensible evidence for choosing the more complex model over the simpler one.
Between the two, BIC penalises extra parameters more heavily than AIC, so if the two measures disagree, BIC is the more conservative choice for a dissertation committee that is wary of overfitting. In this case, both agreed, which made the argument easier to defend.
This is not just a technical formality. Committees want to see that you chose your model because the data demanded it, not because you preferred a particular output.
What Changed In Her Results After Switching Models
Two variables that were statistically insignificant under OLS became significant under negative binomial regression. One variable that looked significant under OLS became insignificant once the model was specified correctly. Her original conclusions were partly wrong, but the corrected version was finally something she could defend without flinching.
In my experience, this pattern is common. Wrong model choice does not just produce wrong numbers, it actively hides or invents relationships that are not really there. I have seen the same thing happen with logistic regression run on continuous outcomes, and with ANOVA run on clearly non-normal data. The mechanism is always the same: the model assumptions and the data do not match, and the output does not warn you about it.
How To Defend A Model Change To Your Dissertation Committee
I never write the analysis chapter for a client. What I do is make sure the student understands the model well enough to defend it under questioning, because that is what actually matters in a viva.
When you switch models mid-dissertation, explain it plainly in your methodology section. State what the original problem was, for example overdispersion or negative predicted values, and explain why the new model resolves it using fit statistics as evidence. If your write-up needs to follow the American Psychological Association’s reporting conventions, report the model type, the offset used, the IRR with its confidence interval, and the fit statistics (AIC, BIC), rather than raw unexponentiated coefficients. Committees are far more comfortable with a corrected model that is explained honestly than a flawed model defended with confidence.
When To Get SPSS Negative Binomial Regression Consulting Help
You do not need help for every regression you run. But certain warning signs suggest you should get a second opinion before your committee finds the problem first.
Signs You Need Help With Count Data Analysis
- Your model produces negative predicted values for a count outcome
- Your committee has questioned model fit or stability
- You are unsure whether to use Poisson, negative binomial, or zero-inflated models
- You have not accounted for exposure or offset variables
- Your significant results keep changing between specifications
What A Consultant Can, And Cannot, Do For You
I can help you identify the right model, check for overdispersion, set up offsets correctly, and interpret your output so you can defend it in your own words, walking into your defence able to explain every result yourself, not just repeat someone else’s write-up.
This distinction matters more than ever. Committees are now scrutinising unusually polished analysis chapters for signs of AI-generated work. A consultant who teaches you the model, rather than handing you a finished chapter, is the safer path either way.
Depending on where you are stuck, this can be a single session to fix one model, or ongoing support through your full analysis chapter. Most students I work with start with one problem, like this negative binomial case, and decide from there whether they need more.
If any of this sounds familiar, my statistical consulting services cover exactly this kind of model troubleshooting for dissertation and thesis researchers.
Frequently Asked Questions
What is negative binomial regression used for?
It is used to model count outcomes, such as number of events or occurrences, when the data shows overdispersion, meaning the variance is larger than the mean.
How do I know if my data is overdispersed?
Check the deviance-to-degrees-of-freedom ratio from your Poisson model. A ratio well above 1 indicates overdispersion, and negative binomial regression is usually a better fit.
What is the difference between Poisson and negative binomial regression?
Poisson assumes the variance equals the mean, while negative binomial adds a dispersion parameter to handle count data where the variance is much larger than the mean.
What is an offset variable, and when do I need one?
An offset accounts for differing exposure across observations, such as population size when modelling event counts across cities of different sizes. You need one whenever your count outcome depends partly on a known exposure that varies across your sample.
How do I interpret an incidence rate ratio (IRR)?
An IRR above 1 means the predictor increases the expected count, an IRR below 1 means it decreases it, and an IRR of 1 means no effect. For example, an IRR of 1.34 means a 34 percent increase in the expected count.
Can I use OLS regression for count data?
Technically SPSS will run it and give you output, but OLS assumes a continuous, normally distributed outcome, so it often produces negative predicted values and unreliable significance for count data.
Should I use R or SPSS for count data regression?
Both can run negative binomial regression correctly. SPSS’s GENLIN interface is easier for students who are not comfortable writing code, while R’s MASS or pscl packages give more flexibility for zero-inflated and hurdle models. The choice usually comes down to what your committee expects and what you are already comfortable troubleshooting.
About the author: I am Siddharth, a statistical consultant and dissertation research coach with over 20 years of analytics consulting experience across R, Python, SPSS, Stata, Power BI, and SQL. I hold an MBA in Finance and an M.Tech, and I run Statssy, helping PhD and thesis researchers get their analysis right and defend it with confidence. Connect with me on LinkedIn.