Skip to content
View demidenm's full-sized avatar

Block or report demidenm

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
demidenm/README.md

👋

Michael Demidenko, PhD

Senior Data Scientist & former Cognitive Neuroscientist

     

Research/Data Science

Current: Senior Data Scientist with expertise in data science, predictive modeling, agentic frameworks, model building and big data engineering. Experienced in survey, finance, behavioral and biological (timeseries) data analysis. Specialized in building scalable data pipelines, measurement assessment, statistical models and automated workflows for 100+ TB datasets.

Former: Cognitive Neuroscientist Researcher at Stanford University.

Technical Skills

Research: Research Design, Experimentation, Data-driven solution, Hypothesis-testing, Survey/Behavioral Analyses, Measurement quality evaluation

Languages: Python, R, Linux/Bash, SQL

Statistical Methods: Exploratory Factor Analysis, Principal Component Analysis, Structural Equation Modeling, Linear/Logistic/Ordinal/Multinomial/Hierarchical Regression, Time Series Analysis, Dimensionality Reduction, A/B Testing, Predictive Modeling, Lasso/Ridge Regression, XGBoost, Random Forest

Python Stack: pandas, numpy, scipy, scikit-learn, statsmodels, matplotlib, seaborn, etc

R Stack: tidyverse, ggplot2, lmer, lavaan, lm/glm, emmeans, psych, etc

Cloud & Infrastructure: AWS (S3, EC2, CLI), Docker, uv, HPC clusters, distributed computing

Data Engineering: ETL pipelines, Git/GitHub, automated workflows, data validation/quality control

Highlighted Projects

PyReliMRI - Python package for statistical reliability analysis in large-scale datasets.

OpenNeuro GLM FitLins - Automated analysis pipeline making 500+ task fMRI datasets more accessible, reducing manual resource costs by 70%+ through containerized simplified downloading/filtering, data reshaping, statistical model building and cloud computing.

ABCD-BIDS E-Prime Processor - Automated workflow for the largest consortium-led study in the United States, processing behavioral data from 20,000+ subjects across 20+ sites. Converts E-Prime files to fMRI-ready format with comprehensive quality control at the subject- and group-level.

HCP-YA Preprocessing - End-to-end processing workflow for behavioral and fMRI data for one of the foundational MRI studies in the US. Processing 1000+ subjects, 28TB dataset, and generating BIDS-compliant descriptives and fitting an HCP and alternative statistical model to the task-based timeseries data.

Professional Highlights

  • Built end-to-end stiatical pipelines processing and analyzing 30+ TB datasets on AWS/HPC
  • Led collaborative teams of researchers, statisticians and analysts
  • Published 35+ peer-reviewed research products
  • Created 5+ open-source packages with openly distribution code code
  • Contributed with data engineering and statistical knowledge to a large-scale ($500M+) NIH-funded study
  • Working with ambiguous problems & highly complex data

For more details, check out my personal webpage

Pinned Loading

  1. stats_ShinyApp stats_ShinyApp Public

    R

  2. PyReliMRI PyReliMRI Public

    Python-based Reliability in MRI: Calculating Group level and Individual Level Similarity between runs and sessions.

    Python 39 3

  3. abcc_datapre abcc_datapre Public

    Python 5

  4. git_traffik git_traffik Public

    Extract & Plot Git Repo Clones/Views Beyond 14-day Insights

    Python

  5. hcpya_preprocess hcpya_preprocess Public

    HCP-YA Preprocessing

    Jupyter Notebook 7

  6. openneuro_glmfitlins openneuro_glmfitlins Public template

    HTML 20 2