Saurabh Ranjan picture

Saurabh Ranjan

  • Cognitive Neuroscience,
  • Neuro AI, Agentic AI,
  • Machine Psychology,
  • Metacognition,
  • Reality Monitoring,
  • Hallucinations.

Ph.D. in Psychology (Behavioral & Cognitive Neuroscience)
University of Florida, 2025.
On The Job Market: Cognitive Science, Neuro AI, AI/ML Research, LLM Safety & Evaluation.

I study how the brain and large language models generate, and sometimes confuse, internal thought with external input, and what that reveals about memory, imagination, and hallucinations in both.

Most work in this space asks whether LLMs think like humans. I probe the inverse of this question, contrasting how humans and LLMs build internal representations, to clarify what is distinct about human cognition. This can not only sharpen theories of human and machine consciousness, but also identify what AI systems need for better self-knowledge — plausibly one prerequisite for agentic AI that can be trusted to self-report and be monitored.

My dissertation, Reality Monitoring in Humans and Artificial Intelligence, found that humans and LLMs share the same core capability that fails during hallucination: telling what you generated yourself from what came from somewhere else. I tested this with matched behavioral experiments in humans and LLM simulations, showing reality monitoring itself generalizes across humans and LLMs as a function of task conditions. I closed the dissertation with a separate finding: the way people's imagined experiences relate to each other is remarkably consistent from person to person, but LLMs couldn't reproduce that pattern at any scale, suggesting a person's internal model of the world isn't just made of individual memories and imagined scenes, but of how those experiences connect to one another.

Generative AI raises the same question at scale. I use the randomized-controlled-trial approach from cognitive psychology to test the general cognitive abilities of large language models. That became Psych Scanner, a framework that helps researchers run cognitive experiments on LLMs at scale instead of writing ad hoc scripts for every study.

I'm currently a Researcher in AI for Biomedical Health Outcomes at UF, co-developing UF's Master's curriculum on Data Science and Agentic LLMs, and building and maintaining the AIBHS Faculty Hub and AI Passport Impact Project sites.

I previously worked with Dr. Brian Odegaard (blue sky, x) at PACLAB, with Dr. Andreas Keil's lab continuing development of Psych Scanner, and with Prof. Narayanan Srinivasan (now at IIT Kanpur) at CBCS.

My publications are on Google Scholar, and my open-source work is on GitHub.

I rarely blog, but when I do, it's on Substack.

Scroll down to read more about highlighted projects and research.

SAURABH RANJAN

Gainesville, FL, U.S.A. • saurabhr.neuroai@proton.mesaurabhr.github.io

GitHubLinkedInBlue SkyXORCID

Ph.D. Cognitive Neuroscientist | Cognitive and Machine Psychology | Human-AI Alignment | NeuroAI


EDUCATION

Ph.D. Psychology: Behavioral & Cognitive Neuroscience University of Florida, USA | 2025

Dissertation: Reality Monitoring in Humans and Artificial Intelligence

M.Sc. Cognitive Science University of Allahabad, India | 2018

Thesis: Intentional Binding in Future-Directed Intentions

B.Sc. Physics Birla Institute of Technology, Mesra, India | 2016

TECHNICAL SKILLS

AI/ML & LLMs: Agentic workflows (LangChain/LangGraph), LLM evaluation & hallucination-detection pipelines, prompt engineering; PyTorch, Hugging Face, scikit-learn, Pydantic, Claude Code (agent testing/tooling)

NeuroAI: Linearizing encoding models, MEG/EEG (MNE), fMRI (Nilearn/SPM), UK Biobank predictive modeling

Stats & Evaluation: Bayesian hierarchical modeling (PyMC, brms), signal detection theory, mixed-effects models (lme4), metacognitive modeling

Data Engineering & MLOps: Python, R, MATLAB, Bash; Pandas, NumPy, Git/GitHub, Docker, SLURM/HPC

Visualization: seaborn, matplotlib, plotly, ggplot2

RESEARCH EXPERIENCE

Researcher: AI for Biomedical Health Outcomes 2026 – Present

University of Florida, Gainesville, FL

• Co-develop UF’s Master’s-level curriculum on Data Science and Agentic LLMs for AI for Biomedical Health Sciences (AIBHS), spanning Foundations, Clinical Application, and Basic Science tracks, part of the program cited in UF’s #1 AI-readiness ranking among large public universities (AIREDEX, 2026); build and maintain the AIBHS Faculty Hub and AI Passport Impact Project sites (20 hands-on modules).

• Built Psych Scanner-Primal, an optimized fork of Psych Scanner for the Prime Intellect Environments Hub, evaluating agentic systems under Reinforcement Learning with Verifiable Rewards (RLVR).

Research Assistant 2025 – 2026

Dr. Andreas Keil’s Lab, University of Florida, Gainesville, FL

• Advanced Psych Scanner from v0.1.0 to v0.4.0, building a harness that automates out-of-distribution (OOD) evaluation of LLMs across 800+ cognitive tasks, giving researchers a reproducible pipeline for running both standard and boutique cognitive paradigms.

Graduate Researcher 2020 – 2025

Dr. Brian Odegaard’s PAC-Lab & UF Department of Psychology, Gainesville, FL

• Discovered that LLM source-attribution accuracy reverses under episodic memory delay, tested across 2 experiments and 6 transformer-based LLMs (Gemma3, Llama3.3, Llama4), a hallucination-relevant failure invisible to single-turn benchmarks.

• Found that corrective feedback produces two distinct metacognitive failure modes, with failure severity tracking active rather than aggregate parameter count, a chain-of-thought-relevant self-knowledge gap that worsens, not improves, with scale (preprint, arXiv:2607.23927).

• Built psychological network models from 2,743 human participants and 6 LLMs (12B–272B parameters): human representational structure replicates across populations (r = 0.31–0.93) while no LLM tested reproduces it at any scale, suggesting a world model can be characterized purely from the structural relationships between internally generated experiences.

• Designed and ran matched experiments with 100+ human participants and parallel LLM simulations to test generalization of reality monitoring across task constraints, complemented by MEG temporal generalization decoding and fMRI encoding models predicting fMRI response from image memorability (NSD dataset), the same seen-vs-imagined question, in humans.

• Authored Psych Scanner (v0.0.1) solo, an open-source framework for running cognitive experiments on LLMs at scale, and built MetaSignal, a Python package for metacognitive/signal-detection analysis.

• Mentored 7 undergraduate researchers and taught a full segment of Physiological Psychology.

Research Assistant 2019 – 2020

Center of Behavioral and Cognitive Sciences, University of Allahabad, India

• Found stronger intentional binding (a measure of sense of agency) for predictive intermediate outcomes than non-predictive ones, replicated across three delays (300/500/700ms) and two contingency levels (OSF preprint).

Research Assistant 2018

Homi Bhabha Centre for Science Education, TIFR, Mumbai, India

• Contributed statistical analysis for Touchy-Feely Vectors (TFV), a gesture-based tool for teaching vectors, across two studies: a 266-student pilot (3 experimental, 3 control classrooms) where the experimental group reasoned differently about vectors and showed higher engagement (IEEE T4E 2019), and a 3-year field study (135 control vs. 131 experimental students) showing improved model-based reasoning in stronger students and engagement in average ones (Journal of Computer Assisted Learning, 2021).

PUBLICATIONS

Working Papers

[1] Ranjan, S., & Odegaard, B. Generalization of generation effect in reality monitoring.

Software

[2] [Preprint] Ranjan, S., Makwana, M., Sokratous, K., & Odegaard, B. (2026). Metasignal: A python package for comprehensive metacognitive analysis and decision-making. arXiv:2607.29093. https://arxiv.org/abs/2607.29093

[1] Ranjan, S., Sokratous, K., & Makwana, M. Psych Scanner: A Framework for Systematic Cognitive Evaluation of Large Language Models.

Theory and Empirical Peer-Reviewed & Preprints

[10] [Preprint] Ranjan, S., Sokratous, K., & Odegaard, B. (2026). Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory. arXiv:2607.23927. ​​https://arxiv.org/abs/2607.23927

[9] [Preprint] Ranjan, S., & Odegaard, B. (2025). Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate. arXiv:2510.04391. https://arxiv.org/abs/2510.04391

[8] Ranjan, S., & Odegaard, B. (2024). Heterarchy or hierarchy? Insights from a new model of visual imagination. Physics of Life Reviews, 49, 74–76.

[7] Ranjan, S., & Odegaard, B. (2024). Reality monitoring and metacognitive judgments in a false-memory paradigm. Neuroscience Research, 201, 3–17.

[6] Maynes, R., Faulkner, R., Callahan, G., Mims, C. E., Ranjan, S., Stalzer, J., & Odegaard, B. (2023). Metacognitive awareness in the sound-induced flash illusion. Philosophical Transactions of the Royal Society B, 378(1886), 20220347.

[5] Chiasson, P., Boylan, M. R., Elhamiasl, M., Pruitt, J. M., Ranjan, S., Riels, K., … & Keil, A. (2023). Effects of neurofeedback training on performance in laboratory tasks: A systematic review. International Journal of Psychophysiology.

[4] Aggarwal, A., & Ranjan, S. (2022). How do undergraduate students reason about ethical and algorithmic decision-making? Proc. 53rd ACM Technical Symposium on Computer Science Education, 488–494.

[3] Karnam, D., Agrawal, H., Parte, P., Ranjan, S., Borar, P., Kurup, P. P., … & Chandrasekharan, S. (2021). Touchy feely vectors: A compensatory design approach to support model-based reasoning. Journal of Computer Assisted Learning, 37(2), 446–474.

[2] Karnam, D., Agrawal, H., Parte, P., Ranjan, S., Sule, A., & Chandrasekharan, S. (2019). Touchy feely affordances of digital technology for embodied interactions can enhance ‘epistemic access’. 2019 IEEE Tenth International Conference on Technology for Education (T4E), 114–121.

[1] [Preprint] Ranjan, S., & Srinivasan, N. (2019). Sense of agency for future-directed intentions. doi.org/10.31234/osf.io/qa93k.

INVITED TALKS

[2] Ranjan, S. (May, 2026). The Spark of Artificial Neuroscience: Psychologically-Grounded Evaluation of Large Language Models. Virtual invited talk, Autonomous Empirical Research Group + Laboratory for Automated Scientific Discovery of Mind and Brain (PI: Dr. Sebastian Musslick), Osnabrück University.

[1] Ranjan, S. (March, 2026). Controlling Imagery Generation and its Awareness. Virtual invited talk, Robert Reinhart Lab, Boston University.

SELECTED CONFERENCE POSTERS & TALKS

[8] Ranjan, S., & Odegaard, B. (2025). Psychological Imagination Networks in Humans and LLMs. Frontiers in NeuroAI, Kempner Institute Symposium, Harvard University.

[7] Ranjan, S., & Odegaard, B. (2025). Visual Imagination Networks in Humans and LLMs. Vision Sciences Society Annual Meeting, St. Pete, FL.

[6] Ranjan, S., & Odegaard, B. (2024). The Fragility of Reality Monitoring under Extraneous Factors. Psychonomic Society 65th Annual Meeting, NYC.

[5] Ranjan, S., & Odegaard, B. (2024). Reality Monitoring, Fast and Slow [Poster + Flash Talk]. Association for Psychological Science, San Francisco. Awarded Scott O. Lilienfeld APS Travel Award.

[4] Maw, M., Zhuang, L., Baltes, J., Ranjan, S., & Odegaard, B. (2024). Task demands and sensory externalization in reality monitoring. PGSO Undergraduate Research Forum, UF. Mentee won 3rd Prize.

[3] Dundigalla, S., Johnson, D., Roh, A., Ranjan, S., & Odegaard, B. (2024). Cognitive strategies in reality monitoring. PGSO Undergraduate Research Forum, UF.

[2] Ranjan, S., Baltes, J., Roh, A., & Odegaard, B. (2023). Confidence in reality monitoring judgments. Journal of Vision, 23(9), 4844.

[1] Ranjan, S., & Srinivasan, N. (2018). Intentional binding and future-directed intentions [Talk]. Annual Conference of Cognitive Science, IIT-Guwahati, India.

AWARDS & FELLOWSHIPS

• 2025: College of Liberal Arts and Sciences Travel Award, University of Florida

• 2025: Threadgill Dissertation Fellowship, University of Florida

• 2024: Scott O. Lilienfeld APS Travel Award, Association for Psychological Science

• 2024–2025: UF Department of Psychology Travel Awards (×4)

• 2020–2025: Graduate Student Fellowship, UF Department of Psychology

• 2016–2018: Graduate Merit Scholarship, Centre of Behavioral & Cognitive Sciences, University of Allahabad

Highlighted Projects

Natural & Artificial Minds

Reality Monitoring in Humans

Raincloud plot showing metacognitive ability (confidence tracking accuracy) is much higher for perceived items than for imagined or new items
  • Built a reality-monitoring task using the DRM (Deese-Roediger-McDermott) false-memory stimulus: participants either perceived or voluntarily imagined the second word in word pairs, then later judged test words as perceived, imagined, or new.
  • Found reality monitoring was better for perceived and new sources than imagined ones, with participants often misremembering imagined items as perceived — an "externalizing bias."
  • Found metacognitive ability (confidence tracking accuracy) was highest for perceived items (gamma = 0.48) versus imagined or new ones (gamma 0.12-0.22), which were statistically indistinguishable from each other.
  • Published as "Reality Monitoring and Metacognitive Judgments in a False-Memory Paradigm," Neuroscience Research (2024).

Reality Monitoring in LLMs

Figure 1 from the preprint: a word-pair paradigm operationalizing reality monitoring in LLMs by contrasting external generation, internal generation, and reality-monitoring source judgments
  • Tested whether six LLMs (Gemma3-12B/27B and their quantization-aware variants, Llama3.3-70B, Llama4-16x17B) show reality monitoring, adapting the same DRM-based paradigm and materials used in the human study.
  • Experiment 1 (single-trial): self-generated items reached ceiling accuracy across all models regardless of memory condition, while externally-provided items showed substantial variability.
Figure 2 from the preprint: three conversational memory architectures (Single-Turn, Trial-Chain, Episodic-Chain) tested across Experiments 1 and 2
  • Experiment 2 (episodic delay): that pattern reversed once a memory delay removed the single-turn shortcut — mean accuracy fell to 66% for perceived items and 48% for imagined items.
  • Found that corrective feedback restructures how source evidence maps to confidence differently across architectures — some models swap their internal/external judgments entirely, others gain accuracy while confidence decouples from correctness — a dissociation invisible to standard accuracy benchmarks and relevant to any AI system operating autonomously across long, multi-turn interactions.
  • Across models, the failure pattern tracked active, not aggregate, parameter count — evidence that evaluating what a model knows isn't enough; tracking where that knowledge came from may matter just as much.
Figure 5 from the preprint: metacognitive sensitivity is consistently higher for external than internal items and is substantially reduced by corrective feedback and larger memory loads
  • Metacognitive sensitivity (confidence tracking accuracy) was consistently higher for externally-provided items than self-generated ones, and corrective feedback sharply reduced it, especially at larger memory loads — models became less able to tell when their own confidence was warranted, not just less accurate.

Human Imagination Psychological Networks

Imagination networks from the PSIQ sensory-modality questionnaire compared across human populations and six LLMs, showing humans cluster consistently while LLMs do not
  • Wrote a discussion piece reframing the heterarchy-vs-hierarchy debate in brain theory through the lens of visual imagination (Physics of Life Reviews, 2024).
  • Built imagination networks from 2,743 participants across Florida, London, and Poland, and compared them against six LLMs (Gemma3-12B/27B, Llama3.3-70B, Llama4-16x17B).
  • Human networks replicated across populations (r = 0.31-0.93); LLMs showed roughly a 5-fold weaker alignment with humans, regardless of model scale (12B-272B parameters) or conversational memory.
  • Most LLM configurations produced degenerate, single-cluster network structure with no human-like organization (median alignment score = 0).
  • Published "Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate" (arXiv:2510.04391, 2025).

Education and Cognition

  • Investigated how undergraduate students reason about ethical and algorithmic decision-making, an early study of human trust in automated decision-making (ACM SIGCSE, 2022).
  • Contributed analysis for "Touchy-Feely Vectors," an embodied digital-geometry tool for grade-11 classrooms, piloted across 266 students (IEEE T4E, 2019).
  • Contributed analysis for a 3-year field study of the tool (135 vs. 131 students), showing it improved model-based reasoning in strong students and engagement in average ones (Journal of Computer Assisted Learning, 2021).
  • Published across IEEE T4E, ACM SIGCSE, and the Journal of Computer Assisted Learning.

Psych Scanner

Psych Scanner wordmark
  • Built a production-grade Python framework (LangChain + LangGraph + Pydantic) for running cognitive experiments with LLMs at scale.
  • Automates and scales "LLMs as a participant" studies, replacing ad hoc scripts with a reusable, tested pipeline.
  • Supports persona management, conversation-history manipulation, and population-level synthesis out of the box.
  • Submitted as a standalone methods paper: "Psych Scanner: A Framework for Systematic Cognitive Evaluation of Large Language Models."
  • Earlier versions of the framework ran the LLM experiments behind two preprints: Reality Monitoring in Large Language Models (arXiv:2607.23927) and Psychological Imagination Networks in Humans and LLMs (arXiv:2510.04391).
  • Welcomes contributions, with full documentation and a contributing guide.

MetaSignal

MetaSignal architecture diagram showing input data flowing through the stdpy SDT core and metacognitive measures into analysis and CLI layers
  • A Python package implementing 20 metacognitive/SDT measures (d′, meta-d′, M-ratio, Type-2 AUC, gamma, and more) plus 6 model-fit diagnostics, all from one function call.
  • Validated against Rahnev (2025)'s MATLAB pipeline across 6 published datasets, replacing the aging MATLAB toolboxes most labs still rely on.
  • Unifies signal detection theory, meta-d′, Bayesian estimation, and information-theoretic metacognition, measures scattered across separate toolboxes, under a single Python library.
  • Open-source, with optional Bayesian (7 estimation approaches) and information-theoretic metacognition modules.
  • Pairs with Psych Scanner to score confidence data from LLM cognitive experiments; the same SDT/metacognition math applies whether the confidence-rated decisions come from a human participant or a model.
  • Welcomes contributions, with full documentation, a contributing guide, and a roadmap.

Software & Technology