Automating Retrieval and Textual Analysis of Scientific Papers

December 16, 2026

Course Description

This workshop teaches you how to build a complete, transparent pipeline for extracting structured knowledge from published scientific papers. You will go from a search query all the way to a visual causal knowledge graph, using only deterministic models, meaning every result is traceable back to an exact sentence in an exact paper.

The workshop is built around a real tool: rasoultilburg/SocioCausaNet, a model specifically trained to detect causal claims in social science text and extract the cause-effect pairs from them

What You Will Learn

  • Bulk paper retrieval: Set up API access and use the eScience Center package to download papers on a topic of your choice.

  • PDF to sentences: Extract text from PDFs, remove page elements and citations, and split the cleaned text into sentences.

  • Causal claims: Use SocioCausaNet to identify sentences that express cause-and-effect relationships and mark the relevant spans.

  • Harmonizing constructs: Explore three approaches for handling constructs that are described differently across papers: a published taxonomy, similarity groups, and clustering.

  • The causal map: Turn causal pairs into a graph, explore causal chains, and trace each connection back to its original sentence.


Prerequisites

This workshop is for:

  • Social science PhD students who want to make their literature reviews faster and more systematic
  • Researchers interested in automating qualitative reviews
  • Anyone who wants to extract and visualize the causal theory embedded in a field’s published literature
  • People curious about NLP tools but with no background in machine learning

You need basic Python familiarity. Nothing else is assumed.


Reading Materials

TBA


Capacity

This course has a maximum capacity of 35 participants.


Time and Location

This workshop will be held on-site only at Eindhoven University of Technologyon December 16, 2026. Details will be provided to all attendees over email after registration for the workshop.

Workshops start from 9:30 to 16:30 with a lunch break from 12:30 to 13:30. Lunch will be provided courtesy of eScience Center.


Registration

This workshop is not yet open for registration.


Instructors

Rasoul Norouzi

Rasoul Norouzi is a PhD candidate in social science at Tilburg University. His research uses natural language processing to support theory development and refinement in the social sciences. He works on language models for causal information extraction from scientific text, graph-based methods for analyzing causal claims, and text mining to support transparent and reproducible research synthesis.

Dr Erik Tjong Kim Sang

Erik Tjong Kim Sang studied Electrical Engineering at the University of Delft and obtained a PhD in computational linguistics at the University of Groningen. Next, he worked as a postdoc and teacher at the universities of Uppsala (Sweden), Antwerp (Belgium), Tilburg, Amsterdam and Groningen. After a postdoc position at the Meertens Institute in Amsterdam, he joined the Netherlands eScience Center as a Senior Research Software Engineer in 2017. Erik has worked on a variety of computational linguistics topics, for example orthographic modeling, syntactic analysis, semantic analysis, named entity recognition, machine translation, summarization, question answering, dialect modeling and analysis of social media text. In his work, he has frequently used machine learning techniques.

Dr Dirk Wulff

Dirk U. Wulff is a senior researcher at the Max Planck Institute for Human Development in Berlin, where he specializes in leveraging big data to explore human decision-making and cognition. He is also the founder of The R Bootcamp and brings extensive experience in helping professionals unlock the power of data storytelling. Dirk is also a Senior Adjunct Researcher at the Center for Cognitive and Decision Science at the University of Basel. His work lies at the intersection of psychology, artificial intelligence, and metascience, and it draws on various methodological approaches, from behavioral experiments to large language mod-els. In addition to his academic work, he is active in data science education for academic and private institutions.