Internships

Models I shipped, data I built, and what changed because of them.

Available for a Summer 2027 internship in ML / LLMs, June to September 2027.

SHEIN

Machine Learning Engineer Intern

Mar – Aug 2026Shanghai, China

Post-trained and shipped a multimodal LLM that screens merchant uploads site-wide for prohibited items.

87% of moderation decisions automated in production
Recall92% → 98%
after GRPO, at flat precision
Duplicates found7.3% → 29.6%
on 50K images, about 4× more
  • Post-trained Qwen3-VL-4B on a 100K chain-of-thought dataset (MS-SWIFT, multi-GPU) and shipped it to production, where it scans merchant uploads for prohibited items.
  • Aligned the SFT checkpoint with RLVR / GRPO, redesigning the accuracy, format and length rewards with asymmetric penalties for misses.
  • Built a three-stage cascaded image deduplication service (MD5 → pHash/dHash → DINOv3 retrieval) with multi-GPU inference and resumable feature caching.
  • Designed a two-layer chain-of-thought auditing tool for hallucinations and answers that contradict their own reasoning. It flagged 7% erroneous CoT in production, filtered out 30% low-quality samples, and raised the L2 model's accuracy to 92%.
Qwen3-VLMS-SWIFTSFTRLVR / GRPODINOv3Multi-GPU

FacePhys

Research Intern

May – Oct 2025Tsinghua University, Beijing

Predicted the full blood-pressure waveform from wearable signals, not just two numbers.

100+ participants in WearNet, the database I built
−30% SBP/DBP error (MAE) vs. a pointwise-loss baseline
  • Built WearNet, a multimodal physiological-signal database that unifies wearable and lab-grade reference recordings with millisecond-level alignment and signal-quality labels.
  • Developed a 1D U-Net that predicts the continuous blood-pressure waveform instead of pointwise systolic and diastolic estimates, giving richer signals for wearable health monitoring.
  • Introduced a combined loss (key-point error, trend deviation, dicrotic-notch constraint) that cut SBP/DBP MAE by 30% on the internal validation set.
PyTorchUNet1DSignal processingWearables

Shanghai AI Laboratory

AI Data Engineering Intern

Mar – May 2025Remote

Turned cultural values into something an LLM can be scored and aligned on.

40+ Chinese cultural values made machine-scorable
+15% alignment score after fine-tuning
−25% unsafe outputs on an internal benchmark
  • Defined a multi-dimensional evaluation framework covering 40+ Chinese cultural values, translating abstract value semantics into structured, machine-scorable metrics for LLM alignment.
  • Curated an SFT / RLHF dataset of 1K value-aligned texts and 800 chosen/rejected preference pairs for fine-tuning.
SFTRLHFPreference dataLLM evaluation

Toolbox

Languages
PythonJavaCC++SQL
Machine learning
PyTorchHugging FaceSFT / RLHF / GRPOMultimodal LLMsRAGMulti-GPU training
Tools
GitLinuxDockerAWSPostgreSQLRedisKubernetesvLLM