About me
Who am I
I am an R programmer, consultant, and research methodologist
specialising in clinical data and evidence synthesis, with expertise in
statistics, reproducible research engineering, and AI-assisted research
automation.
What I do
I build automated, reproducible R pipelines for evidence synthesis
and clinical research, translating research methodology into end-to-end
workflows from data acquisition and processing through analysis to
publication-ready reporting.
Why me
I combine methodological thinking with statistical, programming, and
research expertise. This allows me to approach complex research and data
problems from multiple perspectives and develop solutions that are not
only methodologically sound and technically robust, but also practical,
efficient, and easy to communicate.
What I offer
- Reproducible research pipelines in R: end-to-end
workflows connecting raw data, preprocessing, statistical analysis,
validation, and publication-ready outputs.
- Clinical & real-world data analytics: cleaning,
feature engineering, statistical modelling, diagnostics, and
reproducible reporting for observational and clinical research
data.
- Evidence synthesis methodology & automation:
design research questions, review protocols, extraction frameworks, and
synthesis strategies—including thematic, text-mining, and quantitative
approaches—and translate them into reproducible R and AI-assisted
workflows.
- Statistical methodology: regression modelling,
mixed-effects models, count models, penalised regression, multiple
imputation, psychometrics, and model diagnostics, among others.
- AI-assisted research automation: integration of
language models into screening, document processing, structured
extraction, and other research workflows, with validation and human
oversight.
- Scientific reporting: automated publication-ready
tables, figures, reports, and manuscript components in Quarto and R
Markdown.
- Consulting & training: statistics, R, research
methodology, evidence synthesis, and reproducible analytical
workflows.
At a glance
| Research Automation Developer |
Translate research protocols and analytical plans into reproducible
R pipelines connecting data acquisition, processing, analysis,
validation, and reporting. |
| Clinical & RWE Statistical Programmer /
Analyst |
Build reproducible workflows for clinical and real-world data, from
preprocessing and feature engineering through statistical modelling,
diagnostics, and publication-ready reporting. |
| Evidence Synthesis & Research Methodology
Specialist |
Design review and synthesis methodologies, from research questions,
eligibility criteria, and extraction frameworks through thematic
analysis, text mining, quantitative evaluation, and reproducible
implementation. |
| Research Data Scientist |
Apply advanced statistical methods to clinical, health, survey, and
other empirical research questions while maintaining traceability from
raw data to final results. |
| Reproducible Research Engineer |
Build maintainable research infrastructure using
targets, modular R code, configuration files, dependency
management, testing, logging, and automated reporting. |
| Consultant & Trainer |
Support researchers with statistical methodology, R programming,
research workflows, interpretation, and applied training. |
Independent consultant & small-business owner
Independent consultant who manages research and analytical projects
end-to-end, balancing methodological quality, technical implementation,
communication, and delivery.
Experience Highlights
- Reproducible research engineering: designed
complete analytical and evidence-synthesis systems in R using
targets, modular functions, YAML-based configuration,
renv, automated validation, testing, logging, and Quarto
reporting.
- Clinical & real-world data analytics: processed
messy clinical data into anonymised, analysis-ready datasets; engineered
clinically meaningful variables; fitted statistical models; performed
diagnostics; and generated publication-ready analyses.
- Research methodology & synthesis design:
designed research and evidence-synthesis methodologies from research
questions through analytical strategy, including extraction frameworks,
thematic analysis, text-mining approaches, evaluation of the
psychometric properties of the identified measurement tools, and
integration of qualitative and quantitative evidence.
- Evidence synthesis: built workflows spanning
database searching, bibliographic import, deduplication, title/abstract
screening, full-text assessment, reviewer reconciliation, structured
evidence extraction, and synthesis preparation.
- AI-assisted research workflows: integrated AI into
systematic-review pipelines for screening, document processing,
full-text assessment, and structured extraction while retaining human
review, source-evidence checks, validation, and auditability.
- Advanced applied statistics: applied mixed-effects
logistic regression, negative-binomial regression, bias-reduced logistic
models, restricted cubic splines, penalised regression, multiple
imputation, psychometric validity and reliability analysis, and model
diagnostics.
- Survey & psychometric research: implemented
questionnaire scoring, reverse coding, missing-data rules, scale
reliability assessment, psychometric variables, and downstream
statistical modelling.
- Scientific reporting: automated publication-ready
tables, figures, inline statistics, Word documents, and research reports
so that results remain linked directly to the underlying computational
workflow.
- R development & automation: built reusable
functions and internal R packages for reporting, questionnaire scoring,
file handling, and workflow automation.
- Research consulting: supported 800+ clients across
health, social sciences, economics, and related research domains.
- Training & facilitation: taught practical R
courses for small cohorts (5–12 participants), combining short
theoretical explanations with live coding.
- Project leadership: led small teams and managed
multiple analytical and research projects simultaneously.
- Broader data-science experience: earlier work
includes machine-learning models for gene-expression, socioeconomic, and
sports-related questions.
Skills
Research programming & engineering
- R
- tidyverse
targets / tarchetypes
- Unit testing with
testthat
- Logging and validation
- Reproducible analytical pipelines
- Quarto / R Markdown
renv
- Git / Github
- YAML / configuration-driven workflows
- Functional programming
- Modular R development
- R package development
- Shiny
- SQL (PostgreSQL)
Evidence synthesis
- Systematic-review methodology
- Review protocol operationalisation
- Search strategy implementation
- Bibliographic data processing
- Screening workflows
- Full-text assessment
- Evidence extraction
- Reviewer adjudication
- Narrative-synthesis design and implementation
Statistics & quantitative methods
- Generalised linear models
- Mixed-effects models
- Logistic regression
- Negative-binomial regression
- Penalised regression / LASSO
- Restricted cubic splines
- Multiple imputation
- Rare-event modelling
- Model diagnostics
- Psychometrics and questionnaire scoring
- Descriptive and inferential statistics
- Statistical reporting
- Exploratory machine learning
- Psychometrics measurement literacy
Clinical, health & real-world data
- Clinical-data preprocessing
- Real-world clinical data analytics
- Data anonymisation
- Clinically meaningful feature engineering
- Observational research analytics
- Missing-data workflows
- Survey and questionnaire data
- Data-quality assurance
- Research-data management
- Publication-oriented statistical analysis
AI-assisted research
- LLM-assisted screening
- Structured AI extraction
- Prompt and schema design
- AI-output validation
- Human-in-the-loop workflows
- AI-assisted research automation
Reporting & communication
- Quarto
- R Markdown
ggplot2plots
- Word and PDF reporting
- Publication-ready statistical outputs
- Research writing
- Teaching
- Statistical consulting
- Project management
Selected Projects
AI-assisted systematic-review and evidence-synthesis pipeline
(2025 - 2026)
Role: Lead research methodologist,
evidence-synthesis designer & research automation developer
Research methodology & synthesis design
Served as the primary designer of the review and synthesis
methodology, translating the research questions into a structured
analytical framework for both evidence extraction and downstream
synthesis.
Designed and planned methodological components including:
- conceptualisation of the evidence-synthesis framework,
- thematic analysis of extracted evidence,
- text-mining approaches to identify recurring concepts and patterns
across the literature,
- structured comparison of definitions, dimensions, and
operationalisations of digital maturity,
- psychometric evaluation of measurement instruments identified during
data extraction,
- assessment of how constructs were conceptualised, measured,
validated, and scored across studies,
- extraction structures designed specifically to support subsequent
qualitative and quantitative synthesis.
The methodology was developed together with the computational
workflow rather than treated as a fixed protocol to be implemented
mechanically.
Review workflow
Designed an end-to-end systematic-review workflow covering:
- database-search implementation,
- bibliographic import,
- deduplication,
- title/abstract screening,
- full-text retrieval,
- full-text eligibility assessment,
- PDF and document preprocessing,
- structured evidence extraction,
- reviewer adjudication,
- synthesis preparation,
- quality-control reporting.
AI-assisted research automation
Integrated language models into several stages of the review,
including:
- screening,
- full-text assessment,
- scanned-document processing,
- translation,
- structured evidence extraction.
AI was used as a research tool rather than an autonomous
decision-maker. The workflow therefore incorporated:
- structured JSON schemas,
- model and prompt configuration,
- caching,
- output parsing and repair,
- source-quote validation,
- field-completion checks,
- reviewer-facing inspection reports,
- human-in-the-loop review.
Methodology-to-software translation
Encoded methodological definitions, eligibility criteria, synthesis
dimensions, extraction fields, and analytical requirements in
machine-readable configuration and orchestrated the resulting workflow
using targets and tarchetypes.
This allowed methodological decisions to directly determine how
evidence was searched, screened, extracted, validated, and prepared for
synthesis.
Outcome: a reproducible evidence-synthesis system in
which research methodology, source literature, human decisions,
AI-assisted processing, validation, and downstream synthesis are
connected within a single auditable workflow.
Automated clinical research pipeline for paediatric CT
decision-making (2025 - 2026)
Role: R developer, statistical programmer &
research analyst
Research question & scope
Built an end-to-end R workflow for a clinical research project
evaluating paediatric head-injury CT ordering, guideline concordance,
PECARN-related clinical risk criteria, CT abnormalities, and the effects
of an educational intervention.
The project demonstrates a workflow closely aligned with
real-world clinical data analytics: routinely collected
clinical information is transformed into reproducible evidence
addressing concrete clinical and health-services research questions.
Data engineering
Developed preprocessing workflows for messy spreadsheet-based
clinical data, including:
- import and standardisation,
- anonymisation of patient and physician identifiers,
- recoding and validation,
- clinically meaningful feature engineering,
- codebook generation,
- structured export of analysis-ready datasets.
Statistical methodology
Implemented:
- mixed-effects logistic regression with physician-level random
effects,
- restricted cubic splines for nonlinear age effects,
- likelihood-ratio model comparisons,
- LASSO-based variable screening,
- odds-ratio estimation,
- Haldane-Anscombe correction for sparse cells,
- multiple-testing adjustment,
- simulation-based model diagnostics.
Research engineering & reporting
Organised preprocessing, modelling, diagnostics, and publication
outputs as reproducible pipelines using:
targets,
- central YAML configuration,
renv,
- structured logging,
- unit tests,
- Quarto / R Markdown,
- automated
flextable and ggplot2
outputs.
Outcome: an auditable analysis system in which
clinical transformations, statistical models, diagnostics, tables,
figures, and manuscript results can be regenerated from the underlying
source data.
Reproducible systematic-review workflow for psychosocial factors and
gender identity (2025 - 2026)
Role: Evidence-synthesis researcher & R
workflow developer
Designed R-based infrastructure for a systematic review examining
associations between psychosocial factors and gender-identity constructs
among adults.
Review methodology
Operationalised a review protocol incorporating:
- PEO framing,
- explicit inclusion and exclusion criteria,
- PubMed, Scopus, and Web of Science search strategies,
- controlled vocabulary,
- independent screening by reviewers,
- disagreement resolution,
- full-text eligibility assessment,
- planned risk-of-bias assessment,
- narrative synthesis.
Research infrastructure
Built a targets pipeline covering:
- literature retrieval,
- metadata harmonisation,
- deduplication,
- reviewer screening exports,
- reviewer-decision validation,
- full-text management,
- PDF handling,
- audit logging.
Methodological definitions, operational configuration, raw data,
reviewer decisions, processing logic, and generated outputs were kept
explicitly separated.
Outcome: a transparent and reproducible
evidence-selection process replacing a collection of disconnected manual
review steps with structured research infrastructure.
Reproducible statistical pipeline for vaccination-behaviour research
(2024 - 2025)
Role: Statistical programmer & research
workflow developer
Data & psychometrics
Built a reproducible research workflow for questionnaire-based health
research examining predictors of vaccination delay and refusal.
Implemented:
- raw-data preprocessing,
- demographic recoding,
- psychometric questionnaire scoring,
- reverse scoring,
- item-level missing-data rules,
- scale reliability analysis,
- derived vaccination outcomes.
Missing data & statistical modelling
Designed a multiple-imputation workflow with variable-specific
imputation methods and pooled inference.
Implemented:
- negative-binomial regression for count outcomes,
- bias-reduced logistic regression for vaccine-specific rare
outcomes,
- nested model comparisons,
- adjusted effect estimates,
- incidence-rate ratios and odds ratios.
Scientific quality control
Integrated diagnostics for:
- multicollinearity,
- overdispersion,
- zero inflation,
- influential observations,
- leverage,
- model stability,
- event-per-variable limitations.
Reporting
Connected model outputs directly to publication-ready tables and
manuscript reporting through Quarto and reusable R reporting
functions.
Outcome: a maintainable analysis system spanning raw
survey data, psychometric scoring, missing-data handling, statistical
inference, diagnostics, and publication-ready output.
Cross-national benchmarking of HTA drug-reimbursement decisions
(two-paper series, 2021–2025)
Role: Sole statistician & data
scientist
Data preparation
Started with more than 2,500 reimbursement recommendations provided
as multiple Excel worksheets from 12 health-technology-assessment
agencies.
Developed R workflows to:
- clean, recode, and merge source files,
- repair inconsistent date formats,
- deduplicate entries,
- maintain reproducible updates as new data arrived.
Study design & analytics
Advised the multidisciplinary team on research aims and proposed the
statistical techniques used in the analyses.
Implemented:
- odds-ratio analyses,
- prevalence-adjusted and bias-adjusted kappa (PABAK)
coefficients,
- mixed-effects modelling with drug indication as a random
intercept,
- BCa-bootstrap confidence intervals,
- cross-agency comparisons of time from EMA registration to
recommendation.
Interactive decision support
Built a Shiny dashboard allowing users to compare pairs of HTA
agencies and inspect recommendation profiles and key driver
variables.
Key findings
- Higher clinical added value and lower budget impact were associated
with a higher probability of positive recommendations in many, though
not all, HTA agencies.
- Mixed-effects analysis showed substantial international differences
in time to recommendation, with Wales and Germany generally faster and
Poland showing the longest estimated time to first recommendation.
End-to-end R automation ecosystem for analytics & reporting
(2021–2024)
Role: Sole R-package architect & data
scientist
Scope & build
Designed several interoperable R packages for:
- APA-style statistical reporting,
- psychometric questionnaire scoring,
- Google Drive and Google Sheets file handling,
- client and transaction tracking.
The packages were written in a tidyverse-first style, documented with
roxygen2, and supported by unit testing.
Combined the package ecosystem with R Markdown templates that:
- recognise common statistical object classes,
- insert standardised tables and
ggplot2
visualisations,
- automate common interpretation and reporting steps.
Automation impact
- Reduced analysis-to-report turnaround from days to minutes for more
than 300 studies.
- Reduced copy-paste errors.
- Improved reproducibility and consistency across projects.
Long-term value
The modular package stack became the backbone of the consultancy’s
analytics workflow and was reused across research projects in health,
psychology, social sciences, and other domains.
Additional quantitative projects
Earlier projects demonstrate breadth beyond health research and
evidence synthesis, including:
- structural-transformation analysis of Poland’s hotel sector,
- implementation of Wrocław-taxonomy methods in R,
- socioeconomic predictive modelling,
- gene-expression modelling of clinical endpoints,
- sports analytics,
- interactive Shiny applications.
These projects complement my current specialisation in reproducible
research systems, clinical and real-world data analytics, evidence
synthesis, and research automation.
Publications
- Trzebiński J., Czarnecka J.Z., Cabański M. (2021).
The impact of the narrative mindset on effectiveness in social
problem solving. PLoS ONE, 16(7).
- Trzebiński J., Cabański M., Czarnecka J.Z. (2020).
Reaction to the COVID‑19 pandemic: The influence of meaning in life,
life satisfaction …. Journal of Loss and Trauma,
1‑14.
Education
M.A. Psychology, SWPS University of Social Sciences
and Humanities, 2018
Certificates & Awards
- Rector’s Scholarship for Outstanding Academic Achievement
(2014‑17)
- Toastmasters Competent Leader (2016)
Interests
Scientific reading • Problem solving • Health & wellness •
Philosophy of science • Cognitive psychology • Communication