Skip to content
View mjeans's full-sized avatar

Block or report mjeans

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
mjeans/README.md

Matthew Jeans, PhD

Quantitative research scientist working across nutrition, education, public health, biostatistics, epidemiology, and program evaluation.

With a PhD in Nutritional Sciences, I bring substantive training in nutrition together with experience across education, public health, and applied evaluation. I turn complex administrative, longitudinal, and survey data into evidence that is transparent enough to audit and practical enough to use. My portfolio emphasizes defensible estimands, diagnostic checks, reproducible code, careful uncertainty, and clear boundaries between association, prediction, and causation.

Nutrition, public health, and applied methods

Project Question and methods
NHANES nutrition survey analysis Public two-day dietary recalls, energy-adjusted fiber and sodium density, dietary weights, domain estimates, Taylor-linearized uncertainty, missing-data flow, and survey-weighted regression
Nutrition epidemiology case study Two-day dietary-recall averaging, energy-adjusted fiber and sodium density, completeness reporting, descriptive group contrasts, uncertainty, and measurement-error boundaries
Public Health Methods Lab Nutrition epidemiology, respiratory-disease surveillance, direct age standardization, outbreak risk ratios, rolling signals, Kaplan-Meier analysis, automated tests, and generated outputs
Quasi-experimental program evaluation Propensity-score matching and weighting, common-support and balance diagnostics, clustered inference, regression adjustment, and sensitivity across estimators
Multilevel outcomes analysis Three-level longitudinal modeling, variance decomposition, random effects, contextual variation, interactions, and residual diagnostics
Structural equation modeling Confirmatory factor analysis, measurement invariance, FIML, latent-variable mediation, model diagnostics, and careful noncausal interpretation

Featured nutritional-epidemiology project: The NHANES Nutrition Survey Analysis uses deidentified CDC public-use data to demonstrate two-day dietary measurement, complex-survey inference, missing-data reporting, and descriptive regression.

The Public Health Methods Lab complements it with synthetic, reproducible surveillance, outbreak, nutrition, and time-to-event examples.

Selected nutrition scholarship

My persistent researcher identifier is ORCID 0000-0002-1140-3185. The complete public publication list is available through My NCBI Bibliography, with an additional profile on ResearchGate.

Education, data systems, and responsible analytics

Project What reviewers can inspect
Administrative data pipeline Multisource standardization, deduplication, joins, audit trails, reproducible R/Stata workflows, and validation tests
Evaluation data-quality toolkit Data contracts, domain/range and cross-field rules, issue-level audit output, reusable SQL checks, and CI
Student success predictive modeling Temporal validation, calibration, capacity-aware thresholds, subgroup diagnostics, model cards, and human-review controls
Student success operations dashboard SQL metric layer, dimensional modeling, Power BI-ready measures, implementation monitoring, tested KPIs, and an executive decision memo
SQL analytics case study CTEs, windows, cohorts, anomaly review, tested outputs, metric documentation, and decision-ready interpretation

How I work

  • Start with the decision and estimand. Define the population, comparison, outcome, time window, and interpretation before fitting a model.
  • Make validity visible. Surface missingness, data quality, balance, calibration, clustering, uncertainty, subgroup behavior, and model assumptions.
  • Build for reproduction. Use deterministic synthetic data, executable workflows, tests, continuous integration, data dictionaries, and saved reference outputs.
  • Communicate limits clearly. Separate descriptive, predictive, associational, and causal claims; keep privacy and responsible-use constraints close to the results.

Methods and tools

Nutrition, epidemiology, and biostatistics: dietary recall analysis, energy adjustment, surveillance rates, direct standardization, cohort measures, time-to-event analysis, causal inference, longitudinal and multilevel models, measurement models, missing-data methods, uncertainty and sensitivity analysis
Analysis: R, Stata, Python, SQL
Data and reporting: AWS Athena, Power BI; additional exposure to Tableau, Snowflake, and reproducible Quarto reporting
Delivery: Git, GitHub Actions, automated tests, data contracts, model cards, decision memos, and research governance
Project leadership: research operations, stakeholder engagement, scope and risk management; PMP certification expected August 2026

Portfolio standards

Portfolio projects use either deterministic synthetic records or explicitly documented deidentified public-use data. No client, student, patient, protected health information, restricted records, or row-level public-use files are republished. Each project is designed to expose the full workflow—assumptions, code, quality checks, outputs, interpretation, and limitations—rather than only a polished final chart.

LinkedIn · ORCID · My NCBI Bibliography · ResearchGate

Pinned Loading

  1. student-success-operations-dashboard student-success-operations-dashboard Public

    End-to-end student-success operations analytics with SQL KPIs, a star schema, Power BI-ready measures, data-quality checks, and decision reporting.

    Python

  2. student-success-predictive-modeling student-success-predictive-modeling Public

    Responsible student-success predictive modeling in R with temporal validation, calibration, capacity-aware thresholds, subgroup diagnostics, and reproducible scoring.

    R

  3. administrative-data-pipeline administrative-data-pipeline Public

    Auditable R and Stata pipeline for standardizing, linking, validating, and deduplicating messy multisource administrative data.

    R

  4. quasi-experimental-program-evaluation quasi-experimental-program-evaluation Public

    Reproducible quasi-experimental evaluation using propensity-score matching, balance diagnostics, clustered inference, and robustness checks.

    R

  5. multilevel-outcomes-analysis multilevel-outcomes-analysis Public

    Longitudinal three-level outcomes analysis with mixed-effects models, variance decomposition, diagnostics, and parallel R/Stata implementations.

    R

  6. public-health-methods-lab public-health-methods-lab Public

    Reproducible nutrition epidemiology, surveillance, outbreak, and survival-analysis case studies with synthetic data, tests, and CI.

    Python