Statistical charts reviewed on a laptop

Social Work Vault · Course Resource

Social Statistics

Resource Library
Read-Only

Social Statistics

Statistical foundations for social work: data and measurement, descriptive statistics, probability and inference, hypothesis testing, correlation, regression, program evaluation, and statistical literacy.

100%

Foundations of Social Statistics: Data, Measurement, and Descriptive Statistics

Foundations of Social Statistics: Data, Measurement, and Descriptive Statistics

Statistics is the science of collecting, organizing, analyzing, interpreting, and presenting data. In social work, statistics are essential for understanding social problems, evaluating programs, conducting research, and making evidence-informed decisions. This course provides a comprehensive foundation in statistical reasoning and methods as applied to social work and the social sciences.

Course Overview and Scope

This course equips students with:

  1. An understanding of the role of statistics in social work and social science.
  2. Knowledge of data types, measurement, and sampling.
  3. Skills in descriptive statistics (central tendency, variability, distributions).
  4. Knowledge of graphical and tabular presentation.
  5. Fundamentals of inferential statistics (probability, estimation, hypothesis testing).
  6. Understanding of correlation, regression, and their applications.
  7. Statistical literacy for evaluating research and practice evidence.
  8. Ethical considerations in the use of statistics.

The Role of Statistics in Social Work

Why Social Workers Need Statistics

  • Needs assessment: Quantify community needs (poverty rates, service gaps).
  • Program evaluation: Measure outcomes and effectiveness.
  • Evidence-based practice: Appraise research evidence.
  • Policy analysis: Understand the scale and distribution of social problems.
  • Advocacy: Use data to support social change.
  • Accountability: Document service delivery and outcomes.
  • Clinical practice: Use standardized measures to assess client progress.

The Statistical Method in Social Research

  1. Formulate the research question.
  2. Review existing literature.
  3. Define concepts and variables (operationalization).
  4. Design the study (sampling, data collection).
  5. Collect and organize data.
  6. Analyze data (descriptive and inferential statistics).
  7. Interpret findings.
  8. Disseminate and apply results.
  9. Evaluate and replicate.

Data and Variables

Types of Data

  1. Quantitative data: Numeric data (ages, incomes, test scores, counts).
  2. Qualitative data: Non-numeric data (interviews, written accounts, observations).

Types of Variables

A variable is a characteristic that can take different values.

Levels of Measurement

  1. Nominal: Categories with no order (gender, ethnicity, religious affiliation, type of service). Only frequency counts are meaningful.
  2. Ordinal: Categories with order but unknown distances between points (education level, satisfaction ratings—low/medium/high, social class).
  3. Interval: Equal distances between points, no true zero (temperature in Celsius, IQ scores, test scores). Addition and subtraction are meaningful.
  4. Ratio: Equal distances with a true zero (age, income, number of children, distance). All arithmetic operations are meaningful.

Continuous vs. Discrete Variables

  • Discrete: Can only take certain values (whole numbers—number of children, count of sessions).
  • Continuous: Can take any value within a range (age, height, income, time).

Independent vs. Dependent Variables

  • Independent variable (IV) : The presumed cause or predictor.
  • Dependent variable (DV) : The presumed effect or outcome.

Confounding Variables

Variables that are related to both the IV and DV, potentially explaining the relationship (e.g., poverty affects both housing and child outcomes).

Measurement

Operationalization

Converting abstract concepts into measurable variables. For example, "well-being" might be operationalized as scores on a validated scale (WHO-5, PHQ-9).

Properties of Good Measurement

  1. Validity: The instrument measures what it claims to measure.
  • Content validity: covers the concept.
  • Criterion validity: correlates with an external criterion.
  • Construct validity: measures the underlying construct, consistent with theory.
  1. Reliability: The instrument produces consistent results.
  • Test-retest reliability.
  • Inter-rater reliability.
  • Internal consistency (Cronbach's alpha).
  1. Practicality: Feasible to administer, score, and interpret.

Sampling

Why Sample?

It is usually impractical to study an entire population. Sampling allows inference from a sample to the population.

Key Terms

  • Population: The entire group of interest.
  • Sample: A subset of the population studied.
  • Sampling frame: The list from which the sample is drawn.
  • Sampling error: The difference between sample statistics and population parameters (random variation).
  • Bias: Systematic error that produces unrepresentative samples.

Types of Sampling

Probability Sampling (random selection)

  1. Simple random sampling: Each member has an equal chance of selection.
  2. Systematic sampling: Select every nth member.
  3. Stratified sampling: Divide population into strata, random sample within each.
  4. Cluster sampling: Randomly select clusters (schools, communities), then sample within.
  5. Multi-stage sampling: Combination.

Non-Probability Sampling

  1. Convenience sampling: Easiest to access.
  2. Quota sampling: Sample to match population proportions.
  3. Purposive/judgmental sampling: Select based on criteria.
  4. Snowball sampling: Participants refer others (hard-to-reach populations).

Sample Size

  • Larger samples reduce sampling error.
  • Required size depends on population variability, desired precision, and design.
  • Statistical power analysis determines sample sizes needed to detect effects.

Descriptive Statistics

1. Frequency Distributions

A frequency distribution shows the number (or proportion) of cases in each category or value.

  • For grouped data: class intervals.
  • Relative frequency: Proportions.
  • Cumulative frequency: Running totals.

2. Measures of Central Tendency

  • Mean (arithmetic average) : Sum of values ÷ number of values. Sensitive to outliers.
  • Median: The middle value when data are ordered. Robust to outliers.
  • Mode: The most frequent value. Useful for categorical data.

Choosing a measure:

  • Mean for normally distributed, continuous data.
  • Median for skewed data (income).
  • Mode for nominal data.

3. Measures of Variability/Dispersion

  • Range: Maximum − minimum. Simple but crude.
  • Interquartile range (IQR) : The middle 50% (Q3 − Q1). Robust.
  • Variance: Average squared deviation from the mean.
  • Standard deviation (SD) : Square root of the variance. Expresses dispersion in the data's units.
  • Coefficient of variation: SD ÷ mean (relative variability).
  • Skewness and kurtosis: Describe distribution shape.

4. The Normal Distribution

  • The bell-shaped, symmetric distribution described by mean and standard deviation.
  • Properties: approximately 68% of values within 1 SD of the mean; 95% within 2 SDs; 99.7% within 3 SDs (empirical rule).
  • Many social and biological variables approximate normality.
  • Standard normal distribution (z-scores) : A normal distribution with mean 0 and SD 1. Z-scores express how many SDs a value is from the mean: z = (x − μ) / σ.

5. Descriptive Statistics by Data Type

  • Nominal: frequencies, proportions, mode.
  • Ordinal: frequencies, median, IQR.
  • Interval/ratio: mean, median, SD, variance, range, IQR.

Data Presentation

Tables

  • Frequency tables.
  • Contingency/cross-tabulation tables.

Graphs

  • Bar chart: Categorical data.
  • Histogram: Continuous data distribution.
  • Pie chart: Proportions of a whole (use sparingly).
  • Line graph: Trends over time.
  • Scatterplot: Relationship between two continuous variables.
  • Box plot: Distribution summary (median, quartiles, outliers).
  • Stem-and-leaf plot: Data values + distribution shape.

Principles of Good Data Presentation

  • Clear titles and labels.
  • Appropriate scale (not misleading).
  • Visual simplicity.
  • Honest representation.
  • Accessible to the audience.

Review Questions

  1. Why do social workers need statistics?
  2. Describe the steps of the statistical research method.
  3. Distinguish the four levels of measurement with examples.
  4. Differentiate between independent, dependent, and confounding variables.
  5. Explain validity and reliability in measurement.
  6. Compare probability and non-probability sampling methods.
  7. When should you use the mean versus the median?
  8. Explain the standard deviation and the empirical rule.
  9. Describe the normal distribution and z-scores.
  10. What are the principles of effective data presentation?

Probability and Inferential Statistics: Estimation and Hypothesis Testing

Probability and Inferential Statistics: Estimation and Hypothesis Testing

Inferential statistics allow researchers to draw conclusions about a population from a sample. This module introduces the concepts of probability, sampling distributions, confidence intervals, and hypothesis testing—the foundations of statistical inference in social work research.

Probability

What is Probability?

Probability is the mathematical measure of the likelihood that an event will occur. It ranges from 0 (impossible) to 1 (certain). It can be expressed as a fraction, decimal, or percentage.

Rules of Probability

  1. Complement rule: P(not A) = 1 − P(A).
  2. Addition rule (mutually exclusive events) : P(A or B) = P(A) + P(B).
  3. Addition rule (non-exclusive) : P(A or B) = P(A) + P(B) − P(A and B).
  4. Multiplication rule (independent events) : P(A and B) = P(A) × P(B).
  5. Conditional probability: P(A|B) = P(A and B) / P(B).

The Importance of Probability in Statistics

  • Probability underpins sampling theory.
  • It describes the behavior of samples (sampling distributions).
  • It provides the logic for hypothesis testing and confidence intervals.

Sampling Distributions

The Sampling Distribution of the Mean

If we repeatedly drew samples from a population and computed the mean of each, the distribution of those sample means is the sampling distribution of the mean.

The Central Limit Theorem (CLT)

For sufficiently large samples (typically n ≥ 30), the sampling distribution of the mean is approximately normal, regardless of the population distribution, with:

  • Mean = the population mean (μ).
  • Standard error = σ / √n (the standard deviation of the sampling distribution).

The CLT is central to inference—it justifies using the normal distribution to make inferences about the population.

Standard Error (SE)

The standard error is the standard deviation of the sampling distribution. It decreases as sample size increases. SE = σ / √n.

Estimation

Point Estimates

A point estimate is a single value estimated from the sample (e.g., the sample mean as an estimate of the population mean). Point estimates are imprecise—they do not convey uncertainty.

Interval Estimates / Confidence Intervals (CIs)

A confidence interval provides a range of plausible values for a population parameter, with a specified level of confidence.

The Logic

  • We are 95% confident that the population parameter lies within the interval (mean ± margin of error).
  • The margin of error = critical value × standard error.
  • For 95% confidence, the critical z value is approximately 1.96.
  • CI = x̄ ± z × (σ/√n) for large samples.

Interpretation

  • A 95% CI does NOT mean "95% probability the true value is in this interval" (a common misconception).
  • It means: if we repeated the sampling procedure many times, 95% of the resulting intervals would contain the true population parameter.

Factors Affecting the CI Width

  • Sample size (larger → narrower).
  • Variability in data (larger → wider).
  • Confidence level (higher → wider).
  • Design (e.g., randomization reduces bias).

Hypothesis Testing

What is a Hypothesis?

A hypothesis is a testable statement about a population. The research aims to determine whether the data support the hypothesis.

The Logic of Hypothesis Testing

Hypothesis testing follows the logic of "provisional acceptance" based on probability:

  1. State the null hypothesis (H₀) : A statement of no effect, no difference, or no relationship (e.g., "There is no difference in well-being between program participants and non-participants").
  2. State the alternative hypothesis (H₁ or Ha) : The research hypothesis (e.g., "There is a difference...").
  3. Set the significance level (α) : Usually 0.05 (5%). This is the probability of rejecting the null when it is actually true (Type I error).
  4. Compute the test statistic from the sample data.
  5. Determine the p-value: The probability of obtaining the observed (or more extreme) results if the null hypothesis is true.
  6. Decision:
  • If p < α: reject H₀ (evidence supports the alternative).
  • If p ≥ α: fail to reject H₀ (insufficient evidence).

Statistical Significance vs. Practical Significance

  • Statistical significance: The result is unlikely due to chance (p < α).
  • Practical/clinical significance: The result is meaningful in practice (effect size, real-world impact).
  • Large samples can produce statistically significant but trivially small effects. Always consider effect size.

Types of Errors

| Decision | H₀ True | H₀ False |

|----------|---------|----------|

| Reject H₀ | Type I error (α) : False positive | Correct decision (power) |

| Fail to reject H₀ | Correct decision | Type II error (β) : False negative |

  • Type I error (α) : Rejecting a true null—concluding there is an effect when there is not.
  • Type II error (β) : Failing to reject a false null—missing a real effect.
  • Power = 1 − β: The probability of detecting a real effect (if it exists). High power requires adequate sample sizes.

One-Tailed vs. Two-Tailed Tests

  • Two-tailed: Testing for a difference in either direction (H₁: μ ≠ μ₀). More conservative.
  • One-tailed: Testing for a difference in a specified direction (H₁: μ > μ₀). More powerful but must be pre-specified.

Choosing the Appropriate Statistical Test

| Research Question | IV | DV | Test |

|-------------------|----|----|------|

| Difference in means, 2 groups | Categorical (2) | Continuous | Independent t-test |

| Difference in means, 2 time points | Paired | Continuous | Paired t-test |

| Difference in means, 3+ groups | Categorical | Continuous | ANOVA |

| Association between 2 categorical | Categorical | Categorical | Chi-square |

| Relationship between 2 continuous | Continuous | Continuous | Correlation |

| Prediction of continuous outcome | Continuous/categorical | Continuous | Regression |

| Difference in frequencies | Categorical | Categorical | Chi-square |

Parametric vs. Non-Parametric Tests

Parametric Tests

  • Assume normal distribution and similar variances.
  • More powerful when assumptions are met.
  • Examples: t-tests, ANOVA, Pearson correlation, linear regression.

Non-Parametric Tests

  • Do not assume normality.
  • Used with ordinal data, skewed data, small samples, or when assumptions are violated.
  • Examples: Mann-Whitney U, Wilcoxon test, Kruskal-Wallis, Spearman correlation, chi-square.

p-Values: Interpretation and Misuse

Correct Interpretation

The p-value is the probability of observing the data (or more extreme) if the null hypothesis is true. Small p means the data are unlikely under H₀, supporting H₁.

Common Misinterpretations (to Avoid)

  • "The p-value is the probability the null is true." ✗ (It is not a probability of the hypothesis.)
  • "p = 0.03 means there is a 97% chance the effect is real." ✗
  • "p > 0.05 proves no effect." ✗ (Absence of evidence ≠ evidence of absence.)
  • Dichotomizing results solely on p < 0.05 ignores effect size and confidence intervals.

Best Practices

  • Report effect sizes and confidence intervals, not just p-values.
  • Pre-register hypotheses where appropriate.
  • Interpret findings in context.
  • Avoid "p-hacking" (manipulating analyses until significant).

Review Questions

  1. Define probability and explain its rules.
  2. Explain the Central Limit Theorem and its importance.
  3. What is the standard error and how does sample size affect it?
  4. Explain confidence intervals and their interpretation.
  5. State the null and alternative hypotheses for a study.
  6. Explain Type I and Type II errors and statistical power.
  7. Distinguish one-tailed and two-tailed tests.
  8. Which test would you use to compare means of two independent groups?
  9. Interpret a p-value correctly and identify common misinterpretations.
  10. Why should researchers report effect sizes and confidence intervals?

Tests of Association: Chi-Square, Correlation, and Regression

Tests of Association: Chi-Square, Correlation, and Regression

Much of social work research asks about associations and relationships between variables: Is poverty associated with child maltreatment? Does program participation predict improved well-being? This module explores the major statistical techniques for examining relationships—chi-square, correlation, and regression.

Chi-Square (χ²) Test

Purpose

The chi-square test of independence examines whether two categorical variables are associated (independent vs. dependent).

The Contingency Table

Data are arranged in a cross-tabulation (contingency) table with rows and columns for each variable.

Example: child welfare outcomes by intervention type.

| Outcome | Program A | Program B | Total |

|---------|-----------|-----------|-------|

| Improved | 45 | 30 | 75 |

| Not improved | 15 | 30 | 45 |

| Total | 60 | 60 | 120 |

Calculation

χ² = Σ [(Observed − Expected)² / Expected]

Where Expected = (row total × column total) / grand total for each cell.

Assumptions

  • Independent observations.
  • Sufficient expected frequencies (typically ≥5 in each cell; combine categories if low).
  • Nominal or ordinal data.

Interpretation

  • If χ² is large and p < α: reject H₀—the variables are associated.
  • The chi-square tells you IF there is an association, not its strength.

Measures of Association for Tables

  • Phi (φ) : 2×2 tables.
  • Cramer's V: Larger tables (0 to 1; strength of association).
  • Odds ratio: For 2×2 tables—the odds of the outcome in one group vs. the other.

Goodness-of-Fit Chi-Square

Tests whether observed frequencies fit expected proportions (e.g., equal distribution).

Correlation

Purpose

Correlation measures the strength and direction of a linear relationship between two continuous variables.

The Correlation Coefficient

  • Pearson's r: For continuous, linearly related, roughly normal data.
  • Spearman's rho (ρ) : For ordinal data or skewed/ranked data.
  • Point-biserial: One continuous, one dichotomous variable.

Properties of Pearson's r

  • Ranges from −1 to +1.
  • r = +1: perfect positive relationship.
  • r = −1: perfect negative relationship.
  • r = 0: no linear relationship.
  • Sign: direction. Magnitude: strength.
  • Magnitude interpretation (approximate):
  • |r| < 0.3: weak.
  • 0.3–0.5: moderate.
  • |r| > 0.5: strong.

(Context matters.)

Correlation ≠ Causation

Correlation does not establish causation. A relationship may be due to:

  • Reverse causation.
  • Third/confounding variables.
  • Chance.

Design (experiments, longitudinal, control of confounders) is needed to infer causation.

Correlation Caveats

  • Only measures linear relationships.
  • Outliers can dramatically affect r.
  • Restricted range reduces r.
  • Always inspect the scatterplot.

Simple Linear Regression

Purpose

Regression describes the relationship between a dependent (outcome) variable and one or more independent (predictor) variables, and predicts outcomes.

The Simple Regression Equation

Y = a + bX + e

  • Y: dependent variable (outcome).
  • X: independent variable (predictor).
  • a: intercept (value of Y when X = 0).
  • b: slope (change in Y for a 1-unit change in X).
  • e: error/residual.

Estimation

The line is fitted by least squares: minimizing the sum of squared residuals (vertical distances from the line).

Interpretation Example

If program participation (X = 0, 1) predicts well-being (Y):

  • Slope b = the difference in well-being between participants and non-participants.
  • A significant b (p < .05) supports the program's effect (with appropriate design).

Assumptions

  • Linearity.
  • Independence of observations.
  • Homoscedasticity (constant variance of residuals).
  • Normality of residuals (for inference).
  • No serious outliers.

R² (Coefficient of Determination)

R² is the proportion of variance in the DV explained by the IV(s).

  • R² = 0.25 means 25% of variance is explained. The rest is due to other factors/error.
  • R² ranges from 0 to 1 (higher = better fit, but beware overfitting).

Multiple Regression

Purpose

Multiple regression predicts a continuous outcome from two or more predictors, controlling for each while examining others.

Y = a + b₁X₁ + b₂X₂ + ... + bₖXₖ + e

Why Use Multiple Regression?

  • Control confounding variables: Examine the effect of one predictor while holding others constant.
  • Identify independent predictors.
  • Improve prediction.
  • Test interactions (does the effect of X depend on Z?).

Types

  • Standard/enter: all predictors entered together.
  • Stepwise (forward/backward): automated selection (use with caution—capitalizes on chance).
  • Hierarchical: blocks entered by theory.

Interpretation

  • Unstandardized b: Change in Y per unit of X.
  • Standardized β: Effect in standard deviation units—allows comparison of predictor strength.
  • Significance (p) of each coefficient.
  • for overall model fit.

Categorical Predictors

Categorical predictors are entered as dummy variables (0/1 coding). The reference category is the comparison group.

Logistic Regression

Purpose

Logistic regression predicts a binary outcome (yes/no, event, success) from several predictors.

The Model

Logistic regression models the log-odds (logit) of the outcome: ln[p/(1−p)] = a + bX

Interpretation

  • Odds ratio (OR) : e^b—the multiplicative change in odds for a 1-unit increase in X.
  • OR = 1: no effect.
  • OR > 1: higher odds.
  • OR < 1: lower odds.
  • Example: OR = 1.5 for program participation means participants have 1.5 times the odds of the outcome.

Applications in Social Work

  • Predicting reoffending.
  • Predicting program dropout.
  • Predicting child welfare involvement.
  • Predicting health outcomes.

Choosing and Using Statistical Software

Common Software

  • SPSS: User-friendly menu-driven.
  • R: Powerful, free, open-source, flexible.
  • STATA: Popular in social science research.
  • Excel: Basic descriptive statistics.
  • Python: Data science ecosystem.
  • JASP/JAMOVI: Free point-and-click alternatives to SPSS.

Analytic Workflow

  1. Clean and prepare data.
  2. Explore descriptively (distributions, missing data, outliers).
  3. Check assumptions.
  4. Conduct the analysis.
  5. Interpret and report results.
  6. Consider limitations.

Review Questions

  1. When is the chi-square test appropriate?
  2. How do you calculate and interpret expected frequencies?
  3. Define correlation and interpret Pearson's r.
  4. Why does correlation not imply causation?
  5. Write and interpret the simple regression equation.
  6. Explain R².
  7. Why is multiple regression important for examining relationships?
  8. Explain logistic regression and the odds ratio.
  9. What are the assumptions of linear regression?
  10. Describe the analytic workflow for conducting statistical analysis.

Statistics for Practice Evaluation: Measuring Outcomes and Program Effects

Statistics for Practice Evaluation: Measuring Outcomes and Program Effects

Social workers must evaluate the effectiveness of their interventions. This module focuses on the statistical methods used in program evaluation and single-subject research—measuring outcomes, comparing groups, and demonstrating change.

The Role of Statistics in Program Evaluation

Why Evaluate Programs?

  • Accountability to funders and clients.
  • Evidence-based decision-making.
  • Understanding what works, for whom, and why.
  • Continuous quality improvement.
  • Advocacy for continued funding.
  • Contribution to knowledge.

Types of Evaluation

  1. Needs assessment: What problems exist, and what services are needed?
  2. Process evaluation: Was the program implemented as intended?
  3. Outcome evaluation: Did the program achieve its intended outcomes?
  4. Impact evaluation: Can outcomes be attributed to the program?
  5. Cost-benefit/cost-effectiveness analysis: Is the program worth the cost?

Outcome Measurement

  • Outcome indicators: Observable, measurable changes (symptom reduction, behavior change, service linkage, quality of life).
  • Standardized instruments: Validated measures (PHQ-9 for depression, SDQ for child behavior, WHO-5 for well-being).
  • Satisfaction surveys: Client perceptions.
  • Behavioral measures: Attendance, incidents, completions.
  • Administrative data: Service use, cost data.

Designs for Outcome Evaluation

Experimental Designs (RCT)

  • Participants randomly assigned to intervention or control group.
  • Gold standard for establishing causality.
  • Requires adequate sample size, randomization, and control.
  • Examples: Randomized control trials of parenting programs.

Quasi-Experimental Designs

  • Comparison groups without randomization.
  • Before-after comparisons with a non-equivalent control group.
  • Regression discontinuity: Assignment based on a cutoff score.
  • Interrupted time series: Measure outcomes repeatedly before/after intervention.
  • Strengthen inference but less causal certainty than RCTs.

Pre-Experimental Designs

  • One-group pre-test/post-test: Measure before and after the intervention.
  • Post-test only: Measure only after.
  • Weak designs—no comparison group, so change could result from other factors.

Single-Subject (Single-Case) Designs

Used in clinical and casework practice:

  • Baseline (A) : Repeated measures before intervention.
  • Intervention (B) : Repeated measures during intervention.
  • AB design: Baseline + intervention.
  • ABAB design: Baseline, intervention, withdrawal, re-introduction (stronger evidence).
  • Multiple baseline: Across subjects, behaviors, or settings.
  • Data analyzed visually (level, trend, variability) and statistically (e.g., percentage of non-overlapping data, PAND).

Analyzing Outcome Data

Comparing Groups: t-tests and ANOVA

Independent t-test

  • Purpose: Compare means of two independent groups.
  • Example: Well-being scores of EAP users vs. non-users.
  • Assumptions: normality, homogeneity of variance, independence.

Paired t-test

  • Purpose: Compare means at two time points for the same group.
  • Example: Depression scores before and after intervention (paired observations).

ANOVA (Analysis of Variance)

  • Purpose: Compare means of 3+ groups.
  • Example: Outcomes across three service models.
  • One-way ANOVA: One factor.
  • Two-way ANOVA: Two factors.
  • Post-hoc tests (Tukey, Bonferroni): Which groups differ after a significant ANOVA.
  • ANCOVA: Adjusting for a covariate.

Effect Sizes

Effect size measures the magnitude of the effect, independent of sample size.

  • Cohen's d: (M₁ − M₂) / SD
  • d = 0.2: small; 0.5: medium; 0.8: large.
  • Correlation r: as effect size.
  • Odds ratio: for binary outcomes.
  • Number needed to treat (NNT) : How many people need the treatment for one additional success.

Effect sizes are essential for interpreting practical significance.

Non-Parametric Alternatives

  • Mann-Whitney U (instead of independent t-test).
  • Wilcoxon signed-rank test (instead of paired t-test).
  • Kruskal-Wallis (instead of one-way ANOVA).
  • Used for ordinal/skewed data or assumption violations.

Meta-Analysis and Systematic Reviews

What is a Systematic Review?

A systematic review synthesizes all available evidence on a question using systematic search, appraisal, and synthesis methods.

What is Meta-Analysis?

Meta-analysis statistically combines results from multiple studies to produce an overall effect estimate.

Steps

  1. Formulate the question.
  2. Search comprehensively.
  3. Screen studies.
  4. Appraise quality.
  5. Extract data.
  6. Combine effect sizes (weighted by sample/precision).
  7. Assess heterogeneity.
  8. Report and interpret.

Forest Plot

Displays the effect size (and CI) of each study and the combined estimate.

Heterogeneity

  • I² statistic: proportion of variance due to between-study differences.
  • High heterogeneity challenges pooling.

Single Subject Research in Practice

Applications

  • Evaluating counseling outcomes (pre/post mood ratings).
  • Reducing behaviors (aggressive incidents).
  • Increasing skills (parenting behavior).
  • Measuring client progress over time.

Analysis

  • Visual analysis: Level (mean), trend (slope), variability, immediacy of change, overlap.
  • Statistical methods: Percentage of non-overlapping data points (PND), Tau-U, effect sizes for single-case data.

Ethical Issues in Evaluation

  • Informed consent for participation in research/evaluation.
  • Protection from harm.
  • Privacy and confidentiality of data.
  • Honest reporting (no cherry-picking).
  • Use of data: serving clients, not punishing.
  • Equitable access: Evaluation should not exclude vulnerable groups.
  • Cultural appropriateness of measures.
  • Transparency about methods and limitations.

Review Questions

  1. Explain the types of program evaluation.
  2. What are outcome indicators and how are outcomes measured?
  3. Compare experimental, quasi-experimental, and pre-experimental designs.
  4. Describe single-subject designs and their analysis.
  5. When do you use an independent t-test vs. a paired t-test?
  6. What is ANOVA and when is it used?
  7. Why are effect sizes important and how do you interpret Cohen's d?
  8. What are non-parametric alternatives to t-tests?
  9. Explain systematic review and meta-analysis.
  10. What ethical issues arise in program evaluation?

Statistical Literacy, Ethical Data Use, Case Studies, and Exam Preparation

Statistical Literacy, Ethical Data Use, Case Studies, and Exam Preparation

Statistical literacy—the ability to read, interpret, evaluate, and communicate statistical information—is an essential skill for evidence-based social work practice. This capstone module focuses on critical evaluation of statistics, ethical data use, case studies, and exam preparation.

Statistical Literacy for Social Workers

Reading Research Critically

  • What was the research question?
  • What design was used?
  • How were variables measured (validity, reliability)?
  • What sampling method (bias, representativeness)?
  • What statistical tests? Were assumptions met?
  • Are the results (effect sizes, CIs, p-values) correctly interpreted?
  • What are the limitations?
  • Are the conclusions justified by the evidence?
  • Is it relevant and applicable to practice?

Questions to Ask About a Statistic

  1. Who collected the data and why?
  2. What population does it represent?
  3. How were definitions operationalized?
  4. What is the margin of error?
  5. Are the comparisons fair?
  6. What is the source of funding/interest?
  7. Are effect sizes meaningful?
  8. What stories are hidden by the averages?
  9. Is correlation being presented as causation?

Common Statistical Fallacies in Media and Policy

  • Cherry-picking: Selecting favorable data.
  • Misleading graphs: Truncated axes, inappropriate scales, misleading labels.
  • Confusing significance with importance: Small effects in large samples.
  • Confusing correlation and causation.
  • Ecological fallacy: Inferring individual behavior from group data.
  • Simpson's paradox: Relationships reverse when subgroups are combined.
  • Survivorship bias: Focusing on survivors/existing cases.
  • The "average" as misleading: When distributions are skewed.
  • "No significant difference" misread as "proven equal."

Ethical Use of Statistics

Ethical Principles in Data Use

  1. Honesty: Do not falsify, fabricate, or misrepresent data.
  2. Transparency: Report methods fully; disclose limitations.
  3. Accuracy: Check data quality; avoid errors in analysis and interpretation.
  4. Respect for persons: Protect privacy, obtain consent, honor confidentiality.
  5. Beneficence: Use statistics to promote well-being.
  6. Justice: Avoid bias that harms groups.
  7. Accountability: Take responsibility for analyses and reports.

Data Integrity Issues

  • Plagiarism: Claiming others' work as your own.
  • Fabrication: Making up data.
  • Falsification: Manipulating data/analyses.
  • Data dredging/p-hacking: Examining data until findings are significant.
  • Selective reporting: Reporting only favorable results.
  • Conflict of interest: Financial/personal interests influencing findings.

Ethical Data Collection

  • Informed consent.
  • Voluntary participation.
  • Privacy and confidentiality.
  • Minimal risk.
  • Protection of vulnerable populations (children, elders, incarcerated, mental health clients).

Analyzing Real-World Social Work Data

Case Study 1: Evaluating a Parenting Program

Scenario: A family support agency runs a 12-week parenting program for parents referred by child welfare. To evaluate outcomes, they measure parental stress (PSI) and child behavior (SDQ) before and after the program for 60 parents.

Statistical approach:

  1. Descriptive statistics: Means, SDs, ranges, graphs of both measures.
  2. Paired t-test (or Wilcoxon): Compare pre vs. post scores (same individuals, two time points).
  3. Effect size: Cohen's d for pre-post change.
  4. Confidence intervals for the mean change.
  5. Analysis of subgroups (e.g., by referral type, dosage).
  6. Report: "Parental stress decreased significantly (t(59) = 4.2, p < .001, d = 0.54, 95% CI [2.1, 5.6]), indicating a moderate improvement."

Cautions: No control group → cannot fully attribute change to the program.

Case Study 2: Comparing Two Service Models

Scenario: An agency offers two models of case management (Model A in 3 branches; Model B in 3 others). They compare client outcomes (housing, employment, well-being) after 12 months.

Statistical approach:

  • Chi-square test for categorical outcomes (housing status by model).
  • Independent t-test for continuous outcomes (well-being scores).
  • Control for baseline differences (ANCOVA or matching) if groups differ.
  • Report effect sizes and CIs.

Caution: Branches were not randomly assigned → selection bias risk.

Case Study 3: Identifying Risk Factors

Scenario: A researcher studies risk factors for child protection re-referral among 400 families using case records.

Statistical approach:

  • Descriptive: Frequencies of risk factors, outcomes.
  • Chi-square: Associations between each risk factor and re-referral.
  • Logistic regression: Multiple predictors (poverty, substance abuse, domestic violence, caregiver mental health) predicting re-referral. Report ORs and CIs.
  • Interpretation example: "Caregiver substance abuse was associated with increased odds of re-referral (OR = 2.3, 95% CI [1.4, 3.8]), after controlling for other factors."

Case Study 4: Interpreting Media Statistics

Scenario: A newspaper headline claims "Social Workers Improve Outcomes" based on a study where program participants had 25% better outcomes than non-participants.

Critical questions:

  • Was assignment random? (If not, selection bias.)
  • What were the outcomes and how measured?
  • What is the effect size and CI?
  • Were confounders controlled?
  • Is the claim supported?

Comprehensive Exam Preparation

Multiple Choice Questions

  1. The mean, median, and mode are measures of:

a) Variability

b) Central tendency

c) Association

d) Probability

  1. The standard deviation measures:

a) The average

b) The spread of data around the mean

c) The mode

d) The maximum

  1. A variable measured as "high, medium, low" is:

a) Nominal

b) Ordinal

c) Interval

d) Ratio

  1. The Central Limit Theorem states that for large samples:

a) The population is normal

b) The sampling distribution of the mean is approximately normal

c) The median equals the mean

d) All samples are unbiased

  1. A Type I error is:

a) Accepting a true null

b) Rejecting a true null

c) Accepting a false null

d) Rejecting a false null

  1. Correlation coefficient r ranges from:

a) 0 to 1

b) -1 to +1

c) -100 to 100

d) 0 to 100

  1. Which test compares means of 3+ groups?

a) t-test

b) ANOVA

c) Chi-square

d) Pearson r

  1. A p-value of 0.03 means:

a) 97% chance the null is true

b) The result is practically important

c) If the null is true, the probability of these results is 3%

d) The effect size is large

  1. Logistic regression is used when the outcome is:

a) Continuous

b) Binary

c) Count

d) Time

  1. R² represents:

a) The correlation coefficient

b) The proportion of variance explained

c) The standard error

d) The p-value

Short Answer Questions

  1. Distinguish the four levels of measurement.
  2. Explain the relationship between sample size, standard error, and confidence intervals.
  3. Define Type I and Type II errors and statistical power.
  4. Explain the difference between the mean and the median, and when to use each.
  5. Why does correlation not imply causation?
  6. Describe the chi-square test of independence.
  7. What is an effect size and why is it important?
  8. Name three ethical principles in data use.

Essay Questions

  1. "Statistics are essential for evidence-based social work practice." Discuss the roles of statistics in social work, with examples.
  1. Explain the logic of hypothesis testing, including the null and alternative hypotheses, p-values, and the risks of Type I and Type II errors.
  1. "Correlation does not imply causation." Discuss the relationship between correlation and causation, with examples from social work research.
  1. Describe how you would design and statistically analyze an evaluation of a social work program, including design, measures, tests, and limitations.
  1. "Statistical literacy is a professional competency for social workers." Discuss the skills and ethical responsibilities involved in reading, using, and communicating statistics.

Key Terms Glossary

  • ANOVA: Analysis of variance; compares 3+ group means.
  • Central Limit Theorem: Large-sample sampling distributions are normal.
  • Confidence interval: Range of plausible values for a parameter.
  • Correlation: Measure of linear association.
  • Dependent variable: Outcome variable.
  • Descriptive statistics: Summarizing data.
  • Effect size: Magnitude of effect.
  • Hypothesis testing: Decision procedure for claims about populations.
  • Inferential statistics: Drawing conclusions from samples to populations.
  • Independent variable: Predictor variable.
  • Levels of measurement: Nominal, ordinal, interval, ratio.
  • Mean: Arithmetic average.
  • Median: Middle value.
  • Mode: Most frequent value.
  • Normal distribution: Bell-shaped distribution.
  • p-value: Probability of results if null is true.
  • Regression: Predicting outcomes from predictors.
  • Reliability: Consistency of measurement.
  • Sampling: Selecting a subset of a population.
  • Standard deviation: Spread around the mean.
  • Standard error: SD of the sampling distribution.
  • t-test: Comparing two means.
  • Validity: Accuracy of measurement.
  • z-score: Standardized score in SD units.
This content is protected and read-only. Copying is disabled.