Portrait of Julie (Xinyi) Zhu
Julie (Xinyi) Zhu

Apple Intelligence GenAI Model Evaluation UC Berkeley

Open to GenAI development and ML opportunities in the Bay Area.

Best fit: GenAI model development, LLM applications, multimodal AI, and production ML systems, with evaluation and benchmarking experience to build reliable models.

PROFILE

πŸ‡ΊπŸ‡Έ Work Authorization

US Citizen

πŸ’» Professional Profile

Generative AI Research Engineer specializing in LLM Alignment and Evaluation at Apple Intelligence, with a Master of Information and Data Science from UC Berkeley. Proven track record of driving post-training alignment, model safety, and robustness across 3 WWDC. Expert in designing LLM-as-a-Judge evaluation loops, curating high-quality synthetic data generation pipelines, optimizing model performance, and leveraging adversarial evaluation to secure production-grade, multi-modal workflows at global scale.

πŸš€ Let's Connect

If your company has openings for core Model Development, Fine-Tuning, Model Evaluation tracks, drop me a message! Let's build something incredible. ✨

SKILLS

Generative AI (GenAI)

Agentic AI · Large Language Model (LLM) · Retrieval-Augmented Generation (RAG) · LLM-as-a-Judge · Multimodal AI · Prompt Evaluation · Prompt Engineering · Evaluation Pipelines · Benchmark Design · Tool Calling · Embedding Retrieval · LLM Fine-Tuning · LoRA/QLoRA · LLaMA · Qwen · Gemma · PyTorch · Hugging Face · Transformers · LangChain · Mistral · Cohere

Machine Learning

Natural Language Processing (NLP) · Computer Vision (CV) · Model Evaluation · Model Comparison · Multimodal Evaluation · Benchmarking · Error Analysis · Model Diagnostics · Ranking Systems · ArcFace Embeddings · Logistic Regression · Support Vector Machine (SVM) · Multilayer Perceptron (MLP) · PCA · ROC Analysis · Confusion Matrix · Semantic Similarity · BERT · BLEU · ROUGE · Sentence-Transformer Similarity

Data Science

Python · SQL · Data Engineering · Dataset Curation · Data Cleaning · Feature Engineering · Data Visualization · Data Quality · Performance Analysis · Exploratory Data Analysis (EDA) · Statistical Validation · Tableau · D3.js · PostgreSQL · Altair · Neo4j · R

Software Engineering

Swift · iOS App Development · Java · JavaScript · React · Web Development · HTML · CSS · AJAX · Algorithms · Data Structures · Go · Jinja · MongoDB · AWS · Docker · Git · C++ · C · Assembly Language · Web Scraping

Cultural Fluency

Technical Collaboration in English · Chinese Language and Cultural Grounding · Mandarin Reading, Writing, and Speaking · Simplified Chinese · Traditional Chinese · Bilingual Communication · Cross-Cultural ML Team Collaboration

EXPERIENCE

Apple

Cupertino, CA

June 2022 - Present

Software Development Engineer – LLM Alignment and Evaluation

Apple Intelligence · Multimodal alignment · Safety evaluation · Production monitoring

Multimodal Alignment, Safety, and Evaluation (Image Playground)

  • Developed and executed alignment, correctness, and safety evaluation workflows for Image Playground generative models.
  • Architected specialized LLM-as-a-Judge evaluation loops and metrics to systematically detect failure patterns, stress-test models against adversarial attacks, and enforce strict safety guardrails.
  • Designed and scaled pipelines leveraging synthetic data generation and adversarial evaluation to evaluate text-to-text alignment for core reasoning and international languages, alongside text-to-image alignment, enforcing strict safety guardrails.
  • Audited Chain-of-Thought execution paths to resolve logical inconsistencies and cascading reasoning failures.

End-to-End Model Evaluation and Monitoring

  • Collaborated cross-functionally with machine learning research scientists to translate complex research goals into automated evaluation pipelines, contamination checks, and reproducible variant-comparison workflows for Apple Intelligence features including Image Playground, Visual Intelligence, Math Notes, Smart Script, and Writing Tools.
  • Managed high-throughput parallel execution pipelines for benchmarking and monitoring international datasets across a matrix of Apple Neural Engine (ANE) chips and hardware variants including iPhone, iPad, Mac, and Watch.
  • Constructed rigorous framework-model parity checks to identify output drift, guaranteeing numerical consistency and functional parity between server-side training frameworks and on-device execution environments.
  • Served as the end-to-end evaluation owner, running high-frequency variant comparison reports to programmatically block faulty models while deploying comprehensive Splunk dashboards to isolate cross-team dependencies and bisect regressions.
Apple Park Holiday Party with Craig
Apple Park Holiday Party with Craig
Apple Park campus
Apple Team Retreat
Steve Jobs Theater
Steve Jobs Theater

Feb 2021 - May 2021

Software Development Engineer Intern – Model Evaluation

NLP model testing · multilingual evaluation · asset integrity
  • Engineered dynamic test suites to evaluate and benchmark NLP emoji search models across 41 languages on iPhone and iPad.
  • Implemented automated software testing and triage frameworks to monitor asset delivery and download integrity for NLP models across 47 languages spanning iOS and macOS runtime environments.
Apple Park
Apple Park

Amazon Web Services (AWS)

Seattle, Washington

May 2021 - Aug 2021 · Seattle, Washington

Software Development Engineer Intern

React · JavaScript · Python
  • Developed Apollo website pages using JavaScript, React, HTML, CSS, Python, AJAX, and internal AWS tools.
  • Deployed audit history page and its integration tests to production and finished two extra stretch pages.
Amazon Seattle
Amazon Seattle

Silicon Labs

Austin, Texas

May 2020 - Aug 2020 · Austin, Texas

Product Management Intern

Python · web scraping · competitive analysis
  • Used Python to extract 10 years of Bluetooth data and gathered 30K product details for competitive analysis.
  • Programmed a web-scraping tool to help the company understand growth potential and gain customers.
Silicon Labs Austin
Silicon Labs Austin

UC BERKELEY PROJECTS

Feb 2026 - Apr 2026

Agentic RAG System with LLM-as-a-Judge for Personalized Recommendation

Agentic RAG · LLM-as-a-Judge · Personalized Recommendation · Summarization
  • Built an LLM-powered agent enabling question answering, personalized recommendations, and summarization over course catalog data.
  • Developed a Retrieval-Augmented Generation pipeline with embedding-based retrieval and tool calling to ground responses and execute recommendation workflows.
  • Designed a two-stage ranking system for compliance-aware recommendations using user profiles, learning history, and query signals.
  • Evaluated system performance using LLM-as-a-Judge to assess relevance, grounding, and response quality.
UC Berkeley Capstone
UC Berkeley Capstone

Oct 2025 - Dec 2025

A Proof-of-Concept Evaluation of RAG System

Retrieval-Augmented Generation (RAG) · LangChain · BERT · BLEU · Semantic Similarity · LLM-as-a-Judge
  • Designed and evaluated a Retrieval-Augmented Generation system using LangChain, optimizing embedding models, chunking strategies, and prompt design for distinct engineering and marketing use cases.
  • Implemented a multi-metric evaluation framework with BERT, BLEU, semantic similarity, and LLM-as-a-Judge to assess factual accuracy, semantic alignment, and response quality.
  • Conducted comparative analysis of LLM pipelines, including Mistral vs. Cohere, analyzing trade-offs in accuracy, latency, scalability, and cost to inform deployment recommendations.
UC Berkeley Immersion
UC Berkeley Immersion

Sep 2025 - Dec 2025

Facial Expression Classification

Computer Vision · FER-2013 · ArcFace Embeddings · SVM · Error Analysis
  • Developed an end-to-end facial expression recognition pipeline using the FER-2013 dataset, classifying seven emotions through a combination of traditional computer vision features and deep facial embeddings.
  • Evaluated multiple classifiers, including Logistic Regression, SVM, and MLP, demonstrating that ArcFace embeddings significantly outperformed handcrafted features and achieved up to 63% test accuracy with a nonlinear SVM.
  • Conducted extensive error analysis, PCA-based dimensionality reduction, and multiclass ROC/confusion matrix evaluation to assess generalization, class imbalance, and model trade-offs between accuracy and efficiency.
UC Berkeley Commencement
UC Berkeley Commencement

May 2025 - Aug 2025

BioTitleGen: Title Generator for Medical Abstracts

LLM Fine-Tuning · LoRA/QLoRA · Medical Summarization · ROUGE · Semantic Similarity
  • Developed and fine-tuned large language models, including LLaMA, Qwen, and Gemma, using LoRA/QLoRA for domain-specific medical text summarization.
  • Applied prompt engineering and lightweight fine-tuning to improve title generation accuracy, achieving significant gains in ROUGE and sentence-transformer semantic cosine similarity metrics.
  • Built Python-based pipelines for model training, evaluation, and benchmarking for performance analysis.
UC Berkeley Graduation
UC Berkeley Graduation

Jun 2024 - Aug 2024

World Happiness Report Data Visualization

Tableau · Interactive Dashboards · EDA · Statistical Validation · Data Storytelling
  • Designed and built a multi-dashboard Tableau visualization system on World Happiness Report data, using choropleths, scatter plots, heatmaps, and time-series charts to analyze country-level happiness across 130+ countries, temporal trends, and demographic differences.
  • Built interactive dashboards with filters, parameters, and drill-downs for hypothesis-driven analysis across country, time, and age dimensions, surfacing key trends and relationships in global happiness metrics.
  • Performed exploratory data analysis and statistical validation, including correlation analysis, regression trends, and R2 interpretation, to evaluate feature importance and identify key socio-economic drivers of happiness.
UC Berkeley Class
UC Berkeley Class

EDUCATION

University of California, Berkeley School of Information Master of Information and Data Science

Grade: 3.97/4.00

Aug 2023 - May 2026

Specialized in generative artificial intelligence and machine learning.

  • Generative Artificial Intelligence (GenAI)
  • Natural Language Processing (NLP)
  • Computer Vision (CV)
  • Data Visualization
UC Berkeley campus
UC Berkeley Campus

The University of Texas at Austin Cockrell School of Engineering Bachelor of Science in Electrical and Computer Engineering

Technical GPA: 3.82/4.00

Aug 2018 - May 2022

Specialized in software engineering and data science.

  • Student Leadership Award in Engineering
  • Clinton Sylvester Hartmann Endowed Undergraduate Scholar in Engineering
  • Presidential Scholar
  • Grace Hopper Student Scholar
  • IEEE Computer Society President
Leadership Award
UT Austin Leadership Award

OUTSIDE OF AI

Thanks for scrolling all the way down! Outside of GenAI, here's what keeps me grounded:

πŸ‘©πŸ»β€πŸ« Lifelong Learning & Faith: I spend a lot of time exploring the Bible, reading, and expanding my perspective on how to bring kindness and compassion into everything I do.

πŸ˜‡ Community & Giving Back: I deeply believe in supporting causes that matter. I love volunteering and actively supporting educational, community, and cultural spaces that foster growth and learning.

πŸ‘©πŸ»β€πŸ’» Proud Big Sister: Passion for technology runs in the family! My younger sister is currently studying Computer Science at Carnegie Mellon University and a software engineer intern at Google. I love cheering her on as she enters the tech world.

River of Life Christian Church
River of Life Christian Church
Google Bay View
Google Bay View
Google Family Day
Google Family Day