SHEIN
Machine Learning Engineer Intern
Mar – Aug 2026Shanghai, China
Post-trained and shipped a multimodal LLM that screens merchant uploads site-wide for prohibited items.
87%
of moderation decisions automated in production
Recall92% → 98%
after GRPO, at flat precision
Duplicates found7.3% → 29.6%
on 50K images, about 4× more
- Post-trained Qwen3-VL-4B on a 100K chain-of-thought dataset (MS-SWIFT, multi-GPU) and shipped it to production, where it scans merchant uploads for prohibited items.
- Aligned the SFT checkpoint with RLVR / GRPO, redesigning the accuracy, format and length rewards with asymmetric penalties for misses.
- Built a three-stage cascaded image deduplication service (MD5 → pHash/dHash → DINOv3 retrieval) with multi-GPU inference and resumable feature caching.
- Designed a two-layer chain-of-thought auditing tool for hallucinations and answers that contradict their own reasoning. It flagged 7% erroneous CoT in production, filtered out 30% low-quality samples, and raised the L2 model's accuracy to 92%.
Qwen3-VLMS-SWIFTSFTRLVR / GRPODINOv3Multi-GPU