Limeng Ge (Chloe · 葛力萌)
M.S. in Computer Science, University of Chicago
B.A. in Philosophy, East China Normal University
I'm an M.S. student in Computer Science at the University of Chicago. I started out in philosophy at East China Normal University and moved into machine learning through research and industry work on large language models. Philosophy gave me the questions; engineering gave me ways to test them.
I study how LLMs reason, and where they only pattern-match, how AI is changing the way people write, and how to make what models produce auditable. On the engineering side, I post-train multimodal LLMs with SFT and RL (GRPO) and build research agents that cite their sources and check their own citations.
Do LLMs reason, or do they only sound like they do?
Recent News
- 2026 Project Released DeepTrace, a deep research agent I built from scratch that writes cited reports and checks its own citations. Try the demo →
- 2026 Paper From Output Quality to Auditability: A Data Mining Agenda for LLM-Generated Artifacts accepted at ICDM 2026 BlueSky Track.
- 2026 Life Moved to Chicago to start my M.S. in Computer Science at UChicago.
- 2026 Honor Graduated from East China Normal University as a Shanghai Outstanding Graduate.
- 2026 Work Joined SHEIN as an ML Engineer Intern, post-training a multimodal LLM (SFT + GRPO) for content moderation.
- 2026 Paper Presented Consistent Biases in Large Language Models' Syllogistic Reasoning at the AAAI 2026 Bridge Program on Logical and Symbolic Reasoning in Language Models.
- 2025 Work Research intern at FacePhys (Tsinghua University): built WearNet, a multimodal physiological-signal database with 100+ participants.
- 2025 Honor Received the Creativity Scholarship for Undergraduate Students (top 0.25%) at ECNU.
- 2024 Life Started an exchange year at the College of Arts & Science, New York University.
- 2023 Life Spent two months at a summer school at Oriel College, University of Oxford.
Selected Publications
-
AAAI 2026 Bridge Program on Logical and Symbolic Reasoning in Language ModelsAcross 11,000 syllogisms in all 44 classical forms, five LLMs find the same forms hard (mean pairwise r = 0.886). Every model accepts valid conclusions far more reliably than it rejects invalid ones (a 26–35 point gap), and accuracy rises when the middle term is the grammatical subject.
-
From Output Quality to Auditability: A Data Mining Agenda for LLM-Generated ArtifactsICDM 2026 BlueSky Track Accepted
-
Am I the AI? How AI-Likeness Pressure Reshapes Human Writing ExpressionAAAI 2027 AI for Social Impact Track Under review