Assistant Professor at the Data and Artificial Intelligence Initiative, Faculty of Physical and Mathematical Sciences, University of Chile.
I study how to make natural language AI more interpretable, fair, safe and aligned with human values and rationales.
My research advances Trustworthy Natural Language Processing, with a focus on online harms and high-stakes applications for underrepresented communities in the Global South.
I hold a Ph.D. and an M.Sc. in Natural Language Processing, both in Computer Science and Computational Mathematics from the University of São Paulo
During my doctoral studies, I was a visiting researcher at the University of Southern California, an invited speaker
at the Leibniz Institute for the Social Sciences, and was recognized with the AI for Good award by Bocconi University.
I also actively contribute to the international artificial intelligence and natural language processing research community serving as a program committee member and area chair for venues as
ACL,
NeurIPS,
IJCNN,
AAAI.
I have also co-organized the
International AAAI Conference on Web and Social Media,
Workshop on Online Abuse and Harms,
Workshop on Language Model Interpretability, Safety and Alignment,
Explainable Deep Neural Networks for Responsible AI,
and shared tasks on hate speech [i]
[ii].
I am keen to supervise PhD, MSc, and undergraduate students whose interests align with my research. Feel free to contact me to discuss potential topics of mutual interest.
Research Projects
-
Trustworthy LLMs: Evidence and Rationale Alignment for Hallucination Mitigation
2026: University of Chile -
Equity-Aware Explainable AI for Funding Allocation in Digital Public Goods: Addressing Hate Speech in Latin America
2026: University of Chile -
AI and Global Justice: Interdisciplinary Perspectives on Global Inequalities and Inclusive AI
2026: University of Tübingen -
Building Benchmarks for Hate Speech Detection with Moral Rationales
2025: University of Southern California -
Robust Augmented Retrieval for Natural Language Inference over Transformer-based Models
2025: São Paulo State University & Idiap Research Institute -
Responsible and Explainable Fact-Checking through Fine-Grained Factual Reasoning
2024: Google LARA -
Benchmarking Hate Speech Detection in Hausa Indigenous African Language
2023: University of São Paulo -
Socially Responsible and Explainable Hate Speech Detection in Brazilian Portuguese
2020: University of São Paulo
Research Topics
- Rationale-based and Explainable NLP
- Language Model Interpretability
- Safety and Alignment in LLMs
- Fairness and Bias Mitigation
- Low-Resource Languages
- Content Moderation
- Online Harms