Hi, I’m Paolo! I’m interested in models that keep learning over time and across modalities.
Research interests
Continual learning
Multimodal models
Scaling and efficiency
More about me
I’m a Postdoc at the Artificial Intelligence and Robotics Lab (AIRLab), Politecnico di Milano, working with Prof. Matteo Matteucci. I received my PhD from Politecnico di Milano in 2025, with a thesis on perception and mapping for autonomous driving. In the second half of my PhD, I moved towards continual learning, which is now at the core of my research. In 2024, I was a visiting researcher at Mila - Quebec AI Institute, working on large-scale foundation models with Prof. Irina Rish’s CERC-AAI Lab. I also hold a double MSc degree from Politecnico di Milano and McMaster University (Canada).
I’m always happy to chat about research and possible collaborations: feel free to reach out by email.
News
30 Sep 2026 — Our paper “Towards Characterizing Question Format Bottlenecks for Fine-Tuned Small Language Models” was accepted at the LIGHT workshop at NeurIPS 2026!
Small language models (SLMs) are commonly specialized to a task or domain before deployment, yet the success of that specialization depends partly on how the learned knowledge is queried. We study the performance disparity between binary and multiple-choice (MC) formats after format-specific supervised fine-tuning (SFT), evaluating 15 models from 58M to 3B parameters on binary and multiple-choice QA. On WikiDoc-BMCQA, whose binary and MC questions are generated from the same medical source corpus, closed-book binary verification remains close to chance across much of the SLM range, while four-way MC is learned substantially more strongly relative to its chance baseline. The same qualitative asymmetry appears on public benchmarks, where MC is learned strongly on SciQ and MedMCQA while StrategyQA and the Mintaka yes/no subset show smaller gains over their respective baselines. Supplying the relevant source passage raises binary accuracy to 0.79–0.89 on WikiDoc, showing that verification itself is not intrinsically too hard at this scale when evidence is available. Preliminary diagnostics find that a global yes/no threshold correction yields negligible improvement, while a last-token linear probe exposes no more answer signal than the model’s own output margin. Whether the remaining closed-book bottleneck lies in acquisition, retrieval, or the decision learned by SFT remains open.
@inproceedings{cudrano2026towards,title={Towards Characterizing Question Format Bottlenecks for Fine-Tuned Small Language Models},author={Cudrano, Paolo and Matteucci, Matteo},booktitle={NeurIPS 2026 Workshop on Deployable Small Foundation Models (LIGHT)},year={2026},month=dec,}
We study domain question answering in language models below 100M parameters, a regime where broad world knowledge cannot be stored and specialization to a narrow domain is the only route to competence. In this regime a single design choice—whether a question is posed as balanced yes/no or as multiple choice—turns out to determine whether the model can use its domain knowledge at all, even though a 7B model is indifferent to it. We show this with WikiDoc-BMCQA, a medical benchmark posing identical facts in both formats, with an optional text passage to test closed-book and open-book behavior. Posed as binary yes/no (B), closed-book accuracy of finetuned MobileLLM (125M) and Mamba (130M) falls below a no-knowledge baseline that predicts from surface phrasing alone; posed as multiple choice (MC), the same facts become genuinely accessible above that baseline. Supplying the source passage (open-book) lifts both binary and multiple choice accuracies to almost that of a larger Qwen 7B instruct model, showing that reading abilities are instead saturated even at the 100M scale. The same format and context effects replicate on three disjoint public benchmarks (Mintaka, SciQ). These results suggest a graded, architecture-invariant property of domain knowledge at 100M-scale. While the scale-robust route to using such tiny models remains supplying context, knowledge at this scale is present, but only discriminatively accessible.
@inproceedings{cudrano2026understanding,title={Understanding Without Knowing: Format and Context Impact on 100M-Parameter Domain QA},author={Cudrano, Paolo and Corti, Greta and Pid\'{o}, Sara and Ceccarelli, Valerio and Matteucci, Matteo},booktitle={COLM 2026 Workshop on Methods and Opportunities at Small Scale (MOSS)},year={2026},month=oct,}
Fine-tuning large language models (LLMs) on chain-of-thought (CoT) data shows that a small amount of high-quality data can outperform massive datasets. Yet, what constitutes "quality" remains ill-defined. Existing reasoning methods rely on indirect heuristics such as problem difficulty or trace length, while instruction-tuning has explored a broader range of automated selection strategies, but rarely in the context of reasoning. We propose to define reasoning data quality using influence functions, which measure the causal effect of individual CoT examples on downstream accuracy, and introduce influence-based pruning, which consistently outperforms perplexity and embedding-based baselines on math reasoning within a model family.
@inproceedings{humane2025influence-efficient,title={Influence Functions for Efficient Data Selection in Reasoning},author={Humane, Prateek and Cudrano, Paolo and Kaplan, Daniel Z. and Matteucci, Matteo and Chakraborty, Supriyo and Rish, Irina},booktitle={NeurIPS 2025 Workshops on Efficient Reasoning and Foundations of Reasoning in Language Models},year={2025},month=dec,}
As robotics continues to advance, the need for adaptive and continuously-learning embodied agents increases, particularly in the realm of assistance robotics. Quick adaptability and long-term information retention are essential to operate in dynamic environments typical of humans’ everyday lives. A lifelong learning paradigm is thus required, but it is scarcely addressed by current robotics literature. This study empirically investigates the impact of catastrophic forgetting and the effectiveness of knowledge transfer in neural networks trained continuously in an embodied setting. We focus on the task of visual odometry, which holds primary importance for embodied agents in enabling their self-localization. We experiment on the simple continual scenario of discrete transitions between indoor locations, akin to a robot navigating different apartments. In this regime, we observe initial satisfactory performance with high transferability between environments, followed by a specialization phase where the model prioritizes current environment-specific knowledge at the expense of generalization. Conventional regularization strategies and increased model capacity prove ineffective in mitigating this phenomenon. Rehearsal is instead mildly beneficial but with the addition of a substantial memory cost. Incorporating action information, as commonly done in embodied settings, facilitates quicker convergence but exacerbates specialization, making the model overly reliant on its motion expectations and less adept at correctly interpreting visual cues. These findings emphasize the open challenges of balancing adaptation and memory retention in lifelong robotics and contribute valuable insights into the application of a lifelong paradigm on embodied agents.
@inproceedings{cudrano2024empirical,title={The Empirical Impact of Forgetting and Transfer in Continual Visual Odometry},author={Cudrano, Paolo and Luo, Xiaoyu and Matteucci, Matteo},year={2024},month=jul,booktitle={Third Conference on Lifelong Learning Agents (CoLLAs)}}