ZHANG JR
About
Me
Ph.D. in Computer Science, specializing in post-training for large multimodal models. Currently leading algorithm R&D on VLM/LLM alignment, Agentic RL, and real-time video AI — shipping models that power production-grade intelligent agents.
I believe the next leap in AI is not bigger models, but models that can plan, act, and learn from interaction. My work lives at that frontier: teaching models to use tools, reason through uncertainty, and improve through reinforcement.
When I'm not running training jobs, I'm vibe coding side projects with AI pair programmers, reading new RL papers, or thinking about what it means to build agents that are actually useful.
Senior Algorithm Engineer / Tech Lead · KAOLAURAN Tech Corp
Led algorithm team on full post-training pipeline for large multimodal/language models (VLM/LLM), covering SFT, DPO, On-Policy Distillation, RLVR, and Agentic RL. Built distributed training infrastructure supporting 1.5B–70B+ models and a 15M-sample multimodal data pipeline. Drove deployment of 3 industry-specific LLMs (smart governance, safety production, tobacco), lifting recognition accuracy from 60% to 90%+. Established standardized model iteration workflows, automated multi-dimensional evaluation, and end-to-end training monitoring with 98% data quality pass rate.
Ph.D. Computer Science and Technology · University of Electronic Science and Technology of China
Research focus on multimodal video analysis, including audio-visual speech analysis, video action recognition, and adversarial attacks.
M.S. Communication and Information Systems · Jiangxi University of Science and Technology
Research focus on face recognition techniques including PCA, SVD, manifold learning, and deep neural networks.
B.Eng. Electronic Information Engineering · Wuhan University of Science and Technology
Core coursework in matrix analysis, statistical learning, data structures, and electronic signal analysis and processing.
Technical
Stack
Selected
Work
Video Agent Agentic RL Training System
Replaced rule-based workflow engine with a ReAct-architecture VLM Agent capable of autonomously orchestrating 70+ atomic CV tools. Designed hierarchical action space and multi-dimensional verifiable reward function. Built fully async Inference-Training engine on verl, lifting GPU utilization from 60% to 85%. Composite task completion rate reached 85%+.
Multimodal Reasoning Model
Reproduced and optimized DeepSeek-R1 RLVR framework with a hybrid RL+DPO training strategy. Introduced dynamic importance sampling and clip-higher policy to stabilize training. Implemented On-Policy Distillation enabling cloud-to-edge model evolution.
Real-Time Multimodal Video Interaction System
Designed a low-latency audio-video interaction system supporting visual understanding and external tool calls beyond traditional audio-only voice assistants. Async event-driven pipeline with 20+ processors integrated 15 AI service providers (OpenAI, Google, Anthropic, etc.). Achieved <500ms end-to-end latency supporting 8-channel 1024p@30fps concurrent streams.
Multimodal Training & Data Infrastructure
Built dual-track SFT system for Dense (1.5B–70B) and MoE (~30B-A3B) models on ms-swift, raising GPU utilization to 90%+. Designed modular VLM framework with flexible vision encoder and alignment layer combinations. Constructed end-to-end data pipeline producing 15M multimodal instruction samples at 98% quality pass rate.
Converge — Kubernetes-Inspired Agent Framework
A production-ready Python framework for orchestrating AI agents at scale. Declarative API maps Kubernetes concepts (Deployment→AgentWorkflow, kubelet→AgentRuntime, etcd→StateStore) to agent orchestration — declare the goal, the ReconcileLoop converges to it. Supports DAG-based multi-agent workflows, RBAC tool permissions, Human-in-the-loop for high-risk actions, and full OpenTelemetry observability.
Let's
build
together.