JC
Hi! I'm
Jung In Chang
AI Infrastructure Engineer
Building AI infra where milliseconds matter — from GPU cluster orchestration to inference serving, down to the kernels underneath. Currently investigating how heterogeneous computing shapes LLM inference efficiency on next-generation AI datacenter infrastructure.
{ "name": "Jung In Chang", "location": "San Jose, CA", "role": "AI Infrastructure Engineer", "experience": "3+ years", "education": { "current": "M.S. CS @ UIUC", "past": "B.S. CE @ Boston University" }, "focus": [ "AI Infrastructure", "LLM Inference Serving", "GPU Orchestration" ], "stack": [ "Python", "C++", "Kubernetes", "vLLM / SGLang", "Prometheus" ], "openTo": "Full-time & Internships", "github": "changju784" }
Engineered benchmarking and evaluation workflows for a heterogeneous AI datacenter testbed, measuring LLM inference efficiency across GPU, memory, and network configurations to guide TCO optimization. Established the Kubernetes orchestration foundation for multi-tenant GPU infrastructure, deploying workload isolation and scheduling controls to serve multiple inference frameworks concurrently on shared hardware.
Built and scaled backend infrastructure for maritime SaaS platforms handling high-throughput enterprise workloads. Engineered core services for IMOSX CoCaptain, an automated billing platform with ML-driven laytime calculations, supporting a major enterprise contract. Designed distributed data replication and event-driven pipelines using AWS SNS/SQS and Lambda, improving system latency and horizontal scalability. Streamlined CI/CD pipelines and infrastructure tooling to boost engineering velocity across Veson's global client base.
Developed the core engines of a review analytics platform integrating advanced NLP techniques such as ODP classification, named-entity recognition, and sentiment analysis. Designed and optimized CNN-LSTM architectures, achieving a 10% improvement in F1-score for sentiment prediction. Built an automated web crawler and AWS-based data pipeline to collect and preprocess large-scale text datasets, significantly reducing model training time and manual intervention.
Concentration in LLM RAG and MCP server automation.
Concentration in Machine Learning. Dean's List for 3 semesters.
A multi-tenant Kubernetes simulation that models AI workload lifecycles — Inference, Training, and Data Cleansing — using CPU and RAM as proxies for GPU and VRAM. Demonstrates three core cluster management properties: resource isolation, OOMKill detection, and priority-based scheduling.
Research project measuring how cryptocurrency trading-agent decisions drift when tweet-derived sentiment data is mutated. Uses FinBERT to score BTC/ETH tweets and runs deterministic and FinGPT-style agents on baseline and mutated data windows to quantify divergence under sentiment amplification, temporal jitter, and adversarial tweet injection.
MCP-compatible server exposing automation tools for filing non-emergency 311 service requests in the City of Chicago.
Large-scale Reddit sentiment analysis integrated with market data to forecast short-horizon stock returns via a regularized prediction model.
Fine-tuned vision-language model for weather-aware outfit recommendation using multimodal image and environmental data.
Kotlin Android app automating 5G network tests with real-time Firebase updates. scikit-learn multi-regression model predicting speeds from GPS and altitude data.
Collaborative travel-planning app to create, share, and fork trip itineraries. React + TypeScript frontend with Node.js/Express + MongoDB backend.
Web app using OpenStreet API to find shortest paths via Dijkstra and A* search, with a Python Flask GUI.
Business plan for an automated mooring system replacing manual mooring with sophisticated automation to increase maritime safety and efficiency.
Veson Nautical
CodeCrain Inc.
University of Illinois Urbana-Champaign
Boston University