Analytics dashboard with charts on a laptop screen

Social Work Vault · Course Resource

Program Evaluation

Resource Library
Read-Only

Program Evaluation

Rigorous evaluation frameworks, logic models, qualitative/quantitative metrics, impact assessments, monitoring strategies, and reporting in social programs.

100%

Foundations, Ethics, and Frameworks of Program Evaluation

Foundations, Ethics, and Frameworks of Program Evaluation

Program evaluation is the systematic application of scientific research methods to assess the conceptualization, design, implementation, and utility of social service interventions. Within social work, evaluation serves as the bedrock of evidence-based practice, ethical accountability, and resource stewardship.

1. The Purpose and Taxonomy of Evaluation

```

┌───────────────────────────────────────┐

│ Why Evaluate Social Programs? │

└───────────────────┬───────────────────┘

┌──────────────────────────────┼──────────────────────────────┐

▼ ▼ ▼

┌───────────────┐ ┌───────────────┐ ┌───────────────┐

│ Accountability│ │ Improvement │ │ Knowledge │

│ Funder & donor│ │ Optimize │ │ Contributing │

│ reporting, │ │ service delivery│ │ to evidence │

│ compliance │ │ & workflows │ │ base of field │

└───────────────┘ └───────────────┘ └───────────────┘

```

Major Evaluation Typologies:

  1. Formative Evaluation: Conducted during program planning and early implementation to guide improvements, troubleshoot service delivery, and refine procedures before full rollout.
  2. Summative Evaluation: Conducted at the conclusion of an intervention to judge its overall efficacy, cost-effectiveness, and merit for continuation, scaling, or termination.
  3. Developmental Evaluation (Michael Quinn Patton): Used in complex, rapidly evolving environments where innovations are tested iteratively in real time.

2. The CDC Evaluation Framework

The Centers for Disease Control and Prevention (CDC) Six-Step Evaluation Framework represents the global standard for public health and social programs:

```

[Step 1: Engage Stakeholders]

[Step 2: Describe the Program] (Mission, target population, logic model)

[Step 3: Focus Evaluation Design] (Identify questions, users, methods)

[Step 4: Gather Credible Evidence] (Mixed-methods data collection)

[Step 5: Justify Conclusions] (Analyze data, compare against standards)

[Step 6: Ensure Use & Share Lessons Learned] (Disseminate findings, policy action)

```

3. Program Evaluation Standards (JCSEE)

Evaluations must satisfy four fundamental quality attributes:

  1. Utility: Ensuring the evaluation serves the practical information needs of intended users.
  2. Feasibility: Ensuring procedures are realistic, prudent, diplomatic, and cost-effective.
  3. Propriety: Ensuring ethical conduct, informed consent, respect for participant rights, and protection of vulnerable populations.
  4. Accuracy: Ensuring technical rigor, valid measurement, and sound conclusions.

4. Review Questions

  1. Contrast formative evaluation with summative evaluation, detailing when each should be deployed.
  2. Walk through the six sequential steps of the CDC Program Evaluation Framework.
  3. Explain the four Joint Committee on Standards for Educational Evaluation (JCSEE) standards.

Program Theory, Theory of Change, and Logic Models

Program Theory, Theory of Change, and Logic Models

A program without an explicit theoretical foundation is a "black box" intervention. Clear program modeling illuminates how and why program activities lead to intended outcomes.

1. Theory of Change (ToC) vs. Logic Model

  • Theory of Change (ToC): A high-level strategic roadmap explaining the underlying causal mechanisms, sociological assumptions, and preconditions required to achieve a long-term societal vision.
  • Logic Model: A tactical, systematic graphic illustration of the operational relationships among resources, activities, outputs, and outcomes of a specific program.

2. Anatomy of a Standard Logic Model

```

┌────────────┐ ┌────────────┐ ┌────────────┐ ┌────────────┐ ┌────────────┐

│ INPUTS │──►│ ACTIVITIES │──►│ OUTPUTS │──►│ SHORT-TERM │──►│ LONG-TERM │

│ (Resources)│ │ (Services)│ │ (Products) │ │ OUTCOMES │ │ IMPACTS │

└────────────┘ └────────────┘ └────────────┘ └────────────┘ └────────────┘

Funding Therapy # sessions Knowledge Sustained

Staff Workshops # clients Attitude sobriety

Equipment Job fairs # pamphlets Skills gained Poverty exit

```

| Component | Definition | Concrete Example (Youth Violence Program) |

| :--- | :--- | :--- |

| Inputs | Human, financial, and organizational resources invested | $150,000 grant, 3 social workers, community center space |

| Activities | Specific interventions and processes delivered by the program | Cognitive behavioral anger management workshops, job mentoring |

| Outputs | Direct, quantifiable products of program activities | 50 youth enrolled, 120 workshop hours delivered, 95% attendance |

| Short-Term Outcomes | Immediate changes in learning, knowledge, attitudes, skills (0–6 mos) | 40% increase in conflict de-escalation knowledge on pre/post test |

| Intermediate Outcomes | Observable behavioral or policy changes (6–18 mos) | 60% reduction in school disciplinary suspensions among cohort |

| Long-Term Impact | Ultimate systemic or population-level change (2–5+ yrs) | Decreased youth arrest rates and increased graduation rates in district |

3. Identifying Assumptions and External Risk Factors

  • Assumptions: Underlying beliefs about how the program will operate (e.g., assuming youth have reliable transportation to attend evening sessions).
  • External Factors: Environmental variables outside program control (e.g., local economic recession, changes in juvenile sentencing laws).

4. Review Questions

  1. Distinguish between a Theory of Change and an operational Logic Model.
  2. Differentiate clearly between program Outputs and program Outcomes, providing two examples of each.
  3. Draw a complete logic model for an in-home family preservation program aimed at preventing foster care placement.

Needs Assessment, Feasibility Studies, and Baseline Analysis

Needs Assessment, Feasibility Studies, and Baseline Analysis

Prior to designing and funding social interventions, practitioners conduct rigorous needs assessments to establish genuine service deficits, community assets, and baseline metrics.

1. Conceptual Frameworks of "Need" (Jonathan Bradshaw)

Bradshaw's seminal typology categorizes social need into four dimensions:

  1. Normative Need: Defined by experts, administrators, or statutory benchmarks (e.g., minimum caloric intake, required therapist-to-student ratios).
  2. Felt Need: Subjective desires and perceived wants identified directly by community members (e.g., parents stating they need after-school childcare).
  3. Expressed Need: Felt need translated into concrete action or demand (e.g., long waiting lists for subsidized housing or community mental health clinics).
  4. Comparative Need: Disparities identified by comparing service availability in one geographic community against another similar population.

2. Needs Assessment Methodologies

```

┌───────────────────────────────────┐

│ Needs Assessment Toolset │

└─────────────────┬─────────────────┘

┌────────────────────────────┼────────────────────────────┐

▼ ▼ ▼

┌───────────────┐ ┌───────────────┐ ┌───────────────┐

│ Quantitative │ │ Qualitative │ │ Community │

│ Census data, │ │ Key informant │ │ Asset Mapping │

│ epidemiological│ │ interviews, │ │ Strengths & │

│ health surveys│ │ focus groups │ │ local leaders │

└───────────────┘ └───────────────┘ └───────────────┘

```

  • Secondary Data Analysis: Reviewing hospital discharge data, police crime logs, and census demographic trends.
  • Key Informant Interviews: Qualitative depth interviews with school principals, religious elders, local physicians, and grassroots leaders.
  • Focus Group Discussions (FGDs): Guided group dialogues capturing community norms and nuanced cultural perspectives.
  • Community Asset Mapping: Moving beyond deficit models to map existing strengths, indigenous mutual aid networks, and physical resources.

3. Establishing Baseline Metrics

A baseline is the quantitative or qualitative measurement of the problem before intervention launch. Without a robust baseline measurement, evaluating true program progress or attribution is impossible.

4. Review Questions

  1. Describe Bradshaw's four typologies of social need (normative, felt, expressed, comparative).
  2. Explain the methodology and advantages of Community Asset Mapping over traditional deficit-only needs assessments.
  3. Why is establishing a valid baseline metric critical for subsequent impact evaluations?

Quantitative, Qualitative, and Mixed-Method Evaluation Metrics

Quantitative, Qualitative, and Mixed-Method Evaluation Metrics

Rigorous evaluation demands methodological alignment between evaluation questions and data collection instruments. Mixed-methods designs combine quantitative breadth with qualitative depth.

1. Quantitative Evaluation Metrics & Instrumentation

  • Standardized Psychometric Scales: Validated instruments (e.g., Beck Depression Inventory, PHQ-9, Rosenberg Self-Esteem Scale, Child Abuse Potential Inventory).
  • Measurement Properties:
  • Reliability: Consistency and stability of measurement (Internal consistency Cronbach's $alpha ge 0.80$, Test-retest reliability).
  • Validity: Degree to which an instrument measures what it purports to measure (Construct validity, Content validity, Criterion validity).
  • Administrative Data Tracking: Attendance registers, graduation records, case closure milestones, electronic health record audits.

2. Qualitative Evaluation Methodologies

Capturing the lived experience, nuances of implementation, and unintended consequences:

  • Semi-Structured In-Depth Interviews: Exploring participant journeys, satisfaction, and perceived transformation.
  • Participant Observation: Ethnographic observation of group sessions and agency operations.
  • Document & Thematic Analysis: Inductive coding of case notes and client journals using Braun & Clarke's thematic analysis framework.

3. Mixed-Method Designs (John Creswell)

```

  1. Convergent Parallel Design:

[Quantitative Data Collection & Analysis] ──┐

├──► [Compare & Merge Results]

[Qualitative Data Collection & Analysis] ──┘

  1. Explanatory Sequential Design:

[Quantitative Phase] ──► [Follow-Up Qualitative Phase to Explain Outliers/Mechanisms]

  1. Exploratory Sequential Design:

[Qualitative Exploration] ──► [Develop Instrument] ──► [Quantitative Generalization]

```

4. Review Questions

  1. Distinguish between measurement reliability and measurement validity in program evaluation.
  2. Outline the workflow and purpose of an Explanatory Sequential mixed-methods design.
  3. What strategies can evaluators use to ensure rigor and trustworthiness (credibility, transferability, dependability, confirmability) in qualitative evaluations?

Process, Outcome, and Impact Assessment Strategies

Process, Outcome, and Impact Assessment Strategies

Determining whether observed improvements are genuinely caused by the program requires robust evaluation research designs.

1. Process / Implementation Evaluation

Evaluates fidelity, dosage, and quality of program delivery:

  • Fidelity: Was the intervention delivered exactly as designed in the treatment manual?
  • Dosage: What quantity of the intervention did participants actually receive (hours, sessions)?
  • Reach & Coverage: Did the program reach the intended marginalized sub-groups, or did it experience selective drop-out?

2. Experimental and Quasi-Experimental Research Designs

```

┌───────────────────────────────────┐

│ Evaluation Design Spectrum │

└─────────────────┬─────────────────┘

┌────────────────────────────┼────────────────────────────┐

▼ ▼ ▼

┌───────────────┐ ┌───────────────┐ ┌───────────────┐

│ Experimental │ │ Quasi-Experim.│ │ Non-Experim. │

│ (RCTs) │ │ (Nonequivalent│ │ (Single Group │

│ Random Assign.│ │ Control Group)│ │ Pre/Post Test)│

└───────────────┘ └───────────────┘ └───────────────┘

```

  1. Randomized Controlled Trials (RCTs): Gold standard for causal attribution. Participants randomly assigned to Treatment ($R ightarrow O_1 ightarrow X ightarrow O_2$) or Control ($R ightarrow O_1 ightarrow ext{Control} ightarrow O_2$). Minimizes selection bias.
  2. Quasi-Experimental Designs:
  • Nonequivalent Comparison Group Design: Comparing program participants with a matched non-randomized cohort using Propensity Score Matching (PSM).

- Interrupted Time-Series Design: Multiple observations before and after intervention launch ($O_1, O_2, O_3, O_4 ightarrow X ightarrow O_5, O_6, O_7, O_8$).

  1. Threats to Internal Validity (Donald Campbell & Julian Stanley): History, maturation, testing effects, instrumentation changes, statistical regression to the mean, attrition/mortality, and selection bias.

3. Economic Evaluation: Cost-Benefit vs. Cost-Effectiveness

  • Cost-Effectiveness Analysis (CEA): Compares cost per unit of non-monetary outcome (e.g., $5,000 per child prevented from entering foster care).
  • Cost-Benefit Analysis (CBA): Monetizes both costs and benefits, calculating a Net Present Value (NPV) and Return on Investment (ROI) (e.g., every $1 invested in Nurse-Family Partnership yields $5.70 in societal savings).

4. Review Questions

  1. Explain five major threats to internal validity in non-randomized quasi-experimental program evaluations.
  2. Contrast Cost-Effectiveness Analysis (CEA) with Cost-Benefit Analysis (CBA).
  3. Why is measuring implementation fidelity essential when interpreting poor program outcomes?

Monitoring & Evaluation (M&E) Systems, Data Analytics, and Reporting

Monitoring & Evaluation (M&E) Systems, Data Analytics, and Reporting

Evaluation is incomplete if findings remain buried in unread academic documents. Effective evaluators translate empirical findings into actionable organizational learning, policy reforms, and public dashboards.

1. Establishing an M&E Performance Monitoring System

Continuous performance management tracks routine program indicators in real time:

  • Key Performance Indicators (KPIs): Specific, measurable targets (e.g., % of clients completing case plans within 90 days).
  • Electronic M&E Dashboards: Cloud-based data pipelines aggregating clinical encounters, attendance, and outcome benchmarks for front-line managers.

2. Evaluation Data Visualization and Reporting Standards

```

┌─────────────────────────────────────────────────────────────┐

│ Effective Evaluation Reporting Formats │

├─────────────────────────────────────────────────────────────┤

│ 1. Executive Summary (1-2 pages for executive leadership) │

│ 2. Infographics & Visual One-Pagers (Community audiences) │

│ 3. Comprehensive Technical Report (Funders & researchers) │

│ 4. Stakeholder Data Parties / Interactive Deliberations │

└─────────────────────────────────────────────────────────────┘

```

  • Tailoring to Audiences:
  • Funders: Focus on cost-efficiency, milestone achievement, and ROI.
  • Board of Directors: High-level strategic impacts and risk governance.
  • Clients & Community: Accessible language, visual infographics, and honoring participant contributions.

3. Utilization-Focused Evaluation (UFE)

Championed by Michael Quinn Patton, UFE asserts that evaluations should be judged by their actual utility to primary intended users. Evaluators engage decision-makers throughout the process, preventing evaluations from becoming shelf-ware.

4. Review Questions

  1. Detail the principles of Michael Quinn Patton's Utilization-Focused Evaluation (UFE).
  2. How do routine M&E monitoring systems differ from periodic summative evaluations?
  3. Design a one-page reporting framework for presenting child protection program outcomes to a municipal city council.
This content is protected and read-only. Copying is disabled.