📞 +48 504 532 255   •   ✉️ maciej.statistics@gmail.com   •   linkedin.com/in/caban8/   •   github.com/caban8   •   researchgate.net/profile/Maciej-Cabanski

About me

Who am I

I am an R programmer, consultant, and research methodologist specialising in clinical data and evidence synthesis, with expertise in statistics, reproducible research engineering, and AI-assisted research automation.

What I do

I build automated, reproducible R pipelines for evidence synthesis and clinical research, translating research methodology into end-to-end workflows from data acquisition and processing through analysis to publication-ready reporting.

Why me

I combine methodological thinking with statistical, programming, and research expertise. This allows me to approach complex research and data problems from multiple perspectives and develop solutions that are not only methodologically sound and technically robust, but also practical, efficient, and easy to communicate.

What I offer

  • Reproducible research pipelines in R: end-to-end workflows connecting raw data, preprocessing, statistical analysis, validation, and publication-ready outputs.
  • Clinical & real-world data analytics: cleaning, feature engineering, statistical modelling, diagnostics, and reproducible reporting for observational and clinical research data.
  • Evidence synthesis methodology & automation: design research questions, review protocols, extraction frameworks, and synthesis strategies—including thematic, text-mining, and quantitative approaches—and translate them into reproducible R and AI-assisted workflows.
  • Statistical methodology: regression modelling, mixed-effects models, count models, penalised regression, multiple imputation, psychometrics, and model diagnostics, among others.
  • AI-assisted research automation: integration of language models into screening, document processing, structured extraction, and other research workflows, with validation and human oversight.
  • Scientific reporting: automated publication-ready tables, figures, reports, and manuscript components in Quarto and R Markdown.
  • Consulting & training: statistics, R, research methodology, evidence synthesis, and reproducible analytical workflows.

At a glance

Role Snapshot
Research Automation Developer Translate research protocols and analytical plans into reproducible R pipelines connecting data acquisition, processing, analysis, validation, and reporting.
Clinical & RWE Statistical Programmer / Analyst Build reproducible workflows for clinical and real-world data, from preprocessing and feature engineering through statistical modelling, diagnostics, and publication-ready reporting.
Evidence Synthesis & Research Methodology Specialist Design review and synthesis methodologies, from research questions, eligibility criteria, and extraction frameworks through thematic analysis, text mining, quantitative evaluation, and reproducible implementation.
Research Data Scientist Apply advanced statistical methods to clinical, health, survey, and other empirical research questions while maintaining traceability from raw data to final results.
Reproducible Research Engineer Build maintainable research infrastructure using targets, modular R code, configuration files, dependency management, testing, logging, and automated reporting.
Consultant & Trainer Support researchers with statistical methodology, R programming, research workflows, interpretation, and applied training.

Independent consultant & small-business owner

Independent consultant who manages research and analytical projects end-to-end, balancing methodological quality, technical implementation, communication, and delivery.


Experience Highlights

  • Reproducible research engineering: designed complete analytical and evidence-synthesis systems in R using targets, modular functions, YAML-based configuration, renv, automated validation, testing, logging, and Quarto reporting.
  • Clinical & real-world data analytics: processed messy clinical data into anonymised, analysis-ready datasets; engineered clinically meaningful variables; fitted statistical models; performed diagnostics; and generated publication-ready analyses.
  • Research methodology & synthesis design: designed research and evidence-synthesis methodologies from research questions through analytical strategy, including extraction frameworks, thematic analysis, text-mining approaches, evaluation of the psychometric properties of the identified measurement tools, and integration of qualitative and quantitative evidence.
  • Evidence synthesis: built workflows spanning database searching, bibliographic import, deduplication, title/abstract screening, full-text assessment, reviewer reconciliation, structured evidence extraction, and synthesis preparation.
  • AI-assisted research workflows: integrated AI into systematic-review pipelines for screening, document processing, full-text assessment, and structured extraction while retaining human review, source-evidence checks, validation, and auditability.
  • Advanced applied statistics: applied mixed-effects logistic regression, negative-binomial regression, bias-reduced logistic models, restricted cubic splines, penalised regression, multiple imputation, psychometric validity and reliability analysis, and model diagnostics.
  • Survey & psychometric research: implemented questionnaire scoring, reverse coding, missing-data rules, scale reliability assessment, psychometric variables, and downstream statistical modelling.
  • Scientific reporting: automated publication-ready tables, figures, inline statistics, Word documents, and research reports so that results remain linked directly to the underlying computational workflow.
  • R development & automation: built reusable functions and internal R packages for reporting, questionnaire scoring, file handling, and workflow automation.
  • Research consulting: supported 800+ clients across health, social sciences, economics, and related research domains.
  • Training & facilitation: taught practical R courses for small cohorts (5–12 participants), combining short theoretical explanations with live coding.
  • Project leadership: led small teams and managed multiple analytical and research projects simultaneously.
  • Broader data-science experience: earlier work includes machine-learning models for gene-expression, socioeconomic, and sports-related questions.

Skills

Research programming & engineering

  • R
  • tidyverse
  • targets / tarchetypes
  • Unit testing with testthat
  • Logging and validation
  • Reproducible analytical pipelines
  • Quarto / R Markdown
  • renv
  • Git / Github
  • YAML / configuration-driven workflows
  • Functional programming
  • Modular R development
  • R package development
  • Shiny
  • SQL (PostgreSQL)

Evidence synthesis

  • Systematic-review methodology
  • Review protocol operationalisation
  • Search strategy implementation
  • Bibliographic data processing
  • Screening workflows
  • Full-text assessment
  • Evidence extraction
  • Reviewer adjudication
  • Narrative-synthesis design and implementation

Statistics & quantitative methods

  • Generalised linear models
  • Mixed-effects models
  • Logistic regression
  • Negative-binomial regression
  • Penalised regression / LASSO
  • Restricted cubic splines
  • Multiple imputation
  • Rare-event modelling
  • Model diagnostics
  • Psychometrics and questionnaire scoring
  • Descriptive and inferential statistics
  • Statistical reporting
  • Exploratory machine learning
  • Psychometrics measurement literacy

Clinical, health & real-world data

  • Clinical-data preprocessing
  • Real-world clinical data analytics
  • Data anonymisation
  • Clinically meaningful feature engineering
  • Observational research analytics
  • Missing-data workflows
  • Survey and questionnaire data
  • Data-quality assurance
  • Research-data management
  • Publication-oriented statistical analysis

AI-assisted research

  • LLM-assisted screening
  • Structured AI extraction
  • Prompt and schema design
  • AI-output validation
  • Human-in-the-loop workflows
  • AI-assisted research automation

Reporting & communication

  • Quarto
  • R Markdown
  • ggplot2plots
  • Word and PDF reporting
  • Publication-ready statistical outputs
  • Research writing
  • Teaching
  • Statistical consulting
  • Project management

Selected Projects

AI-assisted systematic-review and evidence-synthesis pipeline (2025 - 2026)

Role: Lead research methodologist, evidence-synthesis designer & research automation developer

Research methodology & synthesis design

Served as the primary designer of the review and synthesis methodology, translating the research questions into a structured analytical framework for both evidence extraction and downstream synthesis.

Designed and planned methodological components including:

  • conceptualisation of the evidence-synthesis framework,
  • thematic analysis of extracted evidence,
  • text-mining approaches to identify recurring concepts and patterns across the literature,
  • structured comparison of definitions, dimensions, and operationalisations of digital maturity,
  • psychometric evaluation of measurement instruments identified during data extraction,
  • assessment of how constructs were conceptualised, measured, validated, and scored across studies,
  • extraction structures designed specifically to support subsequent qualitative and quantitative synthesis.

The methodology was developed together with the computational workflow rather than treated as a fixed protocol to be implemented mechanically.

Review workflow

Designed an end-to-end systematic-review workflow covering:

  • database-search implementation,
  • bibliographic import,
  • deduplication,
  • title/abstract screening,
  • full-text retrieval,
  • full-text eligibility assessment,
  • PDF and document preprocessing,
  • structured evidence extraction,
  • reviewer adjudication,
  • synthesis preparation,
  • quality-control reporting.

AI-assisted research automation

Integrated language models into several stages of the review, including:

  • screening,
  • full-text assessment,
  • scanned-document processing,
  • translation,
  • structured evidence extraction.

AI was used as a research tool rather than an autonomous decision-maker. The workflow therefore incorporated:

  • structured JSON schemas,
  • model and prompt configuration,
  • caching,
  • output parsing and repair,
  • source-quote validation,
  • field-completion checks,
  • reviewer-facing inspection reports,
  • human-in-the-loop review.

Methodology-to-software translation

Encoded methodological definitions, eligibility criteria, synthesis dimensions, extraction fields, and analytical requirements in machine-readable configuration and orchestrated the resulting workflow using targets and tarchetypes.

This allowed methodological decisions to directly determine how evidence was searched, screened, extracted, validated, and prepared for synthesis.

Outcome: a reproducible evidence-synthesis system in which research methodology, source literature, human decisions, AI-assisted processing, validation, and downstream synthesis are connected within a single auditable workflow.


Automated clinical research pipeline for paediatric CT decision-making (2025 - 2026)

Role: R developer, statistical programmer & research analyst

Research question & scope

Built an end-to-end R workflow for a clinical research project evaluating paediatric head-injury CT ordering, guideline concordance, PECARN-related clinical risk criteria, CT abnormalities, and the effects of an educational intervention.

The project demonstrates a workflow closely aligned with real-world clinical data analytics: routinely collected clinical information is transformed into reproducible evidence addressing concrete clinical and health-services research questions.

Data engineering

Developed preprocessing workflows for messy spreadsheet-based clinical data, including:

  • import and standardisation,
  • anonymisation of patient and physician identifiers,
  • recoding and validation,
  • clinically meaningful feature engineering,
  • codebook generation,
  • structured export of analysis-ready datasets.

Statistical methodology

Implemented:

  • mixed-effects logistic regression with physician-level random effects,
  • restricted cubic splines for nonlinear age effects,
  • likelihood-ratio model comparisons,
  • LASSO-based variable screening,
  • odds-ratio estimation,
  • Haldane-Anscombe correction for sparse cells,
  • multiple-testing adjustment,
  • simulation-based model diagnostics.

Research engineering & reporting

Organised preprocessing, modelling, diagnostics, and publication outputs as reproducible pipelines using:

  • targets,
  • central YAML configuration,
  • renv,
  • structured logging,
  • unit tests,
  • Quarto / R Markdown,
  • automated flextable and ggplot2 outputs.

Outcome: an auditable analysis system in which clinical transformations, statistical models, diagnostics, tables, figures, and manuscript results can be regenerated from the underlying source data.


Reproducible systematic-review workflow for psychosocial factors and gender identity (2025 - 2026)

Role: Evidence-synthesis researcher & R workflow developer

Designed R-based infrastructure for a systematic review examining associations between psychosocial factors and gender-identity constructs among adults.

Review methodology

Operationalised a review protocol incorporating:

  • PEO framing,
  • explicit inclusion and exclusion criteria,
  • PubMed, Scopus, and Web of Science search strategies,
  • controlled vocabulary,
  • independent screening by reviewers,
  • disagreement resolution,
  • full-text eligibility assessment,
  • planned risk-of-bias assessment,
  • narrative synthesis.

Research infrastructure

Built a targets pipeline covering:

  • literature retrieval,
  • metadata harmonisation,
  • deduplication,
  • reviewer screening exports,
  • reviewer-decision validation,
  • full-text management,
  • PDF handling,
  • audit logging.

Methodological definitions, operational configuration, raw data, reviewer decisions, processing logic, and generated outputs were kept explicitly separated.

Outcome: a transparent and reproducible evidence-selection process replacing a collection of disconnected manual review steps with structured research infrastructure.


Reproducible statistical pipeline for vaccination-behaviour research (2024 - 2025)

Role: Statistical programmer & research workflow developer

Data & psychometrics

Built a reproducible research workflow for questionnaire-based health research examining predictors of vaccination delay and refusal.

Implemented:

  • raw-data preprocessing,
  • demographic recoding,
  • psychometric questionnaire scoring,
  • reverse scoring,
  • item-level missing-data rules,
  • scale reliability analysis,
  • derived vaccination outcomes.

Missing data & statistical modelling

Designed a multiple-imputation workflow with variable-specific imputation methods and pooled inference.

Implemented:

  • negative-binomial regression for count outcomes,
  • bias-reduced logistic regression for vaccine-specific rare outcomes,
  • nested model comparisons,
  • adjusted effect estimates,
  • incidence-rate ratios and odds ratios.

Scientific quality control

Integrated diagnostics for:

  • multicollinearity,
  • overdispersion,
  • zero inflation,
  • influential observations,
  • leverage,
  • model stability,
  • event-per-variable limitations.

Reporting

Connected model outputs directly to publication-ready tables and manuscript reporting through Quarto and reusable R reporting functions.

Outcome: a maintainable analysis system spanning raw survey data, psychometric scoring, missing-data handling, statistical inference, diagnostics, and publication-ready output.


Cross-national benchmarking of HTA drug-reimbursement decisions (two-paper series, 2021–2025)

Role: Sole statistician & data scientist

Data preparation

Started with more than 2,500 reimbursement recommendations provided as multiple Excel worksheets from 12 health-technology-assessment agencies.

Developed R workflows to:

  • clean, recode, and merge source files,
  • repair inconsistent date formats,
  • deduplicate entries,
  • maintain reproducible updates as new data arrived.

Study design & analytics

Advised the multidisciplinary team on research aims and proposed the statistical techniques used in the analyses.

Implemented:

  • odds-ratio analyses,
  • prevalence-adjusted and bias-adjusted kappa (PABAK) coefficients,
  • mixed-effects modelling with drug indication as a random intercept,
  • BCa-bootstrap confidence intervals,
  • cross-agency comparisons of time from EMA registration to recommendation.

Interactive decision support

Built a Shiny dashboard allowing users to compare pairs of HTA agencies and inspect recommendation profiles and key driver variables.

Key findings

  • Higher clinical added value and lower budget impact were associated with a higher probability of positive recommendations in many, though not all, HTA agencies.
  • Mixed-effects analysis showed substantial international differences in time to recommendation, with Wales and Germany generally faster and Poland showing the longest estimated time to first recommendation.

End-to-end R automation ecosystem for analytics & reporting (2021–2024)

Role: Sole R-package architect & data scientist

Scope & build

Designed several interoperable R packages for:

  • APA-style statistical reporting,
  • psychometric questionnaire scoring,
  • Google Drive and Google Sheets file handling,
  • client and transaction tracking.

The packages were written in a tidyverse-first style, documented with roxygen2, and supported by unit testing.

Combined the package ecosystem with R Markdown templates that:

  • recognise common statistical object classes,
  • insert standardised tables and ggplot2 visualisations,
  • automate common interpretation and reporting steps.

Automation impact

  • Reduced analysis-to-report turnaround from days to minutes for more than 300 studies.
  • Reduced copy-paste errors.
  • Improved reproducibility and consistency across projects.

Long-term value

The modular package stack became the backbone of the consultancy’s analytics workflow and was reused across research projects in health, psychology, social sciences, and other domains.


Additional quantitative projects

Earlier projects demonstrate breadth beyond health research and evidence synthesis, including:

  • structural-transformation analysis of Poland’s hotel sector,
  • implementation of Wrocław-taxonomy methods in R,
  • socioeconomic predictive modelling,
  • gene-expression modelling of clinical endpoints,
  • sports analytics,
  • interactive Shiny applications.

These projects complement my current specialisation in reproducible research systems, clinical and real-world data analytics, evidence synthesis, and research automation.


Publications

  • Trzebiński J., Czarnecka J.Z., Cabański M. (2021). The impact of the narrative mindset on effectiveness in social problem solving. PLoS ONE, 16(7).
  • Trzebiński J., Cabański M., Czarnecka J.Z. (2020). Reaction to the COVID‑19 pandemic: The influence of meaning in life, life satisfaction …. Journal of Loss and Trauma, 1‑14.

Education

M.A. Psychology, SWPS University of Social Sciences and Humanities, 2018


Certificates & Awards

  • Rector’s Scholarship for Outstanding Academic Achievement (2014‑17)
  • Toastmasters Competent Leader (2016)

Interests

Scientific reading • Problem solving • Health & wellness • Philosophy of science • Cognitive psychology • Communication