Report2026
Ziwei Fan

Ziwei Fan
ML Research Engineer at Apple · Ph.D. in Computer Science, University of Illinois Chicago
I build intelligent systems for personalization & recommendation and code migration & translation — including personalized ranking models, multi-modal LLM fine-tuning, agentic test generation, and repo-scale code documentation.
01 Research
Google Scholar →My research centers on personalized generation & retrieval and large language models (LLMs) for code. I design personalized ranking models that address data sparsity, representation collapse, and cross-domain shift in recommender systems, and I build LLM agents that migrate and translate code at repository scale while aligning multi-modal LLMs with real-world tail signals. More recently, I explore proactive, just-in-time agents that infer when to act and re-rank content on-device, without an explicit request.
Proactive Just-in-Time Agents
On-device agents that decide when to act, what to ask, and how to re-rank — as in Kairosrank.
Personalized Ranking Models
Sequential and graph-based ranking for large-scale recommendation and user modeling.
LLM for Code
Self-debugging, agentic test generation, and evaluation for code translation & migration — shipping in AWS Transform for Mainframe.
Representation & Diversity
Mitigating representation collapse; controllable diversification for recommendation.
02 News
03 Experience
ML Research Engineer
Apple
Applied Scientist II
Amazon
- Prime Video Homepage & Notifications (RecSys). Homepage ranking model product delivery with positive stream-hour lift. Bayesian day-of-week / hour-of-day optimization for in-app notifications with additional WW stream-hour gain.
- AWS Transform for Mainframe — Mainframe / Java code migration agent (CodeLLM). KDD ’24
- Agentic Automatic Test Generation for system-level functional-equivalence testing with differential fuzzing; condition/exception execution-path extraction via AST + LLM improves over strong random-testing baselines.
- Agentic Code Documentation for long (repo-level, 20K+ LoC per file) context extraction with test-time context optimization for hallucination reduction.
- LLM-as-Judge Ensemble for quality measurement of generated documents via perturbations and claim-level metric definitions.
- LLM post-training for prompt compression via adaptive meta-token fine-tuning. arXiv ’25arXiv ’26
- Self-debug+ agent with format verifier — measurable success-rate@1 improvement for Java Upgrade.
- AWS Clean Rooms ML — Audience Expansion (RecSys). Transformer user-sequence embeddings and seed-user expansion for ads campaigns. Launched at re:Invent 2023.
- Multi-modal LLM fine-tuning. Implicit attribute-value extraction (EIVEN) with tail-attribute improvements. NAACL ’24
- Personalized recommendation explanation. Aspect-instructed, logic-scaffolded explanation generation with LLMs. WSDM ’24
Research / Applied Scientist / Data Scientist Interns
Salesforce Research · AWS AI · Spotify Research · Stitch Fix
- Product knowledge-graph pre-training for zero-shot item-based recommendation (Salesforce Research). CIKM ’23
- Personalized federated domain adaptation for item-to-item recommendation (AWS AI). UAI ’23
- Discovery-episode ranking via multi-source augmentations (Spotify Research). arXiv ’23
- Data science for personalized styling (Stitch Fix).
Graduate Research & Teaching Assistant
University of Illinois Chicago
- Transformer encoders for sequential recommendation. WWW ’22CIKM ’21BigData ’22
- Data augmentation & denoising for long-tail recommendation. SIGIR ’21SIGIR ’23
- Representation collapse, diversity, and robustness. WWW ’23arXiv ’24
- Federated / transfer learning for cross-domain recommendation. KDD ’24
04 Education
Ph.D., Computer Science
University of Illinois Chicago
Aug 2018 — May 2023
Thesis: Data Sparsity in Recommender Systems — Collaborative Neighbors Enrichment. Advised by Philip S. Yu.
M.S., Computer Science
Purdue University — Indianapolis (IUPUI)
Aug 2016 — May 2018
B.Eng., Network Engineering
South China Agricultural University
Sep 2012 — Jun 2016
05 Selected Publications
Full list on Google Scholar →Report2026