The ideal candidate can effectively bridge AI model requirements ↔ hardware capabilities ↔ customer expectations, guiding customers from model selection → hardware sizing → deployment decisions → production readiness.

What You’ll Do

AI Model Porting & Optimization Deploy, optimize and scale deep learning AI models onto accelerator‑based data center platforms, including: Model conversion workflows Quantization techniques (INT8 / mixed precision) Runtime integration and optimization Integrate ML models onto Qualcomm’s Cloud AI ML stack from frameworks such as PyTorch, TensorFlow, and ONNX. Drive improvements in model throughput, latency, and accuracy, with clear trade‑off analysis. Build, test, and deploy scalable inference pipelines using serving frameworks such as vLLM, TGI, and Triton. Optimize workloads for LLM and GenAI models across both multi-SoC and multi-card architectures. Collaborate with engineering teams to analyze and refine training and inference for advanced deep learning applications. Identify bottlenecks across compute, memory, and runtime, and guide optimization strategies. Contribute to Qualcomm’s Cloud AI GitHub repository and developer documentation, sharing technical best practices and solutions. Develop and integrate end-to-end ML application pipelines with customer frameworks and libraries. Customer‑Facing Technical Engagement Act as a trusted technical advisor for customers deploying AI workloads. Engage in hardware sizing and architecture discussions, aligning model requirements with infrastructure capabilities. Provide technical guidance on: AI model selection Deployment feasibility System architecture and performance expectations Lead discussions on model capabilities and limitations based on real customer use cases. Model–Infrastructure Alignment Assess and evaluate AI model requirements and recommend alternative model approaches when necessary. Align model characteristics (latency, throughput, accuracy) with accelerator and system capabilities. Connect model requirements with: Memory constraints Accelerator architecture Scaling limitations Support customers in defining model selection strategies based on deployment realities. Performance & Scalability Engineering Evaluate performance characteristics of AI models in production scenarios, including: Throughput expectations Latency targets Concurrency behavior Guide architecture decisions around: Scaling strategies (horizontal vs vertical) Hardware deployment sizing Contribute to discussions on: Workload scalability limits Impact of model selection on system performance and efficiency Provide insights into capacity planning and infrastructure optimization. End‑to‑End AI Pipeline Design Drive discussions around end‑to‑end AI pipelines, including: Multi‑model workflows (e.g., detection + tracking + recognition) Data preprocessing and post‑processing stages Guide decisions on video and data processing stacks, including: Video pipeline choices (e.g., FFMPEG vs GStreamer) Integration into inference pipelines Ensure pipelines are aligned with: Performance requirements Hardware capabilities Real‑time constraints Model Trade‑off Analysis & Validation Highlight and explain trade‑offs between: Accuracy vs compatibility Model quality vs deployment feasibility Support decision‑making on: Model simplification vs performance gains Precision vs efficiency trade‑offs Lead or support model capability validation in deployment environments. Collaborate with customers to define: Inference assumptions Model sizing strategies for large‑scale workloads

Required Qualifications

Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience). 10–15+ years of experience in: Deep learning model development or deployment experience on CPUs/GPUs/ASICs. Inference systems and optimization Data center or edge AI platforms Strong experience with: Model quantization and optimization techniques AI model frameworks (e.g., PyTorch, TensorFlow) Model deployment pipelines Excellent C/C++/Python programming and software design skills, including debugging, and performance analysis. Hands on expertise with Linux-based systems, low level software, drivers, and system bring up. Proven ability to analyze and optimize model performance in production environments. Solid understanding of: AI inference hardware constraints System level performance bottlenecks Strong communication skills and experience in customer facing technical roles. Willingness to travel for customer engagements and strategic reviews.

Preferred Qualifications

Skilled in deploying models on platforms that use hardware accelerators for inference. Experienced with managing multi-model workflows and building real-time AI systems, including computer vision, video, and analytics projects. Knowledgeable about distributed inference methods and handling large-scale model deployments. Proficient in developing and maintaining video processing workflows and using relevant software frameworks. Deep understanding of how system-level decisions affect performance in actual deployment environments. Capable of simplifying complex technical ideas into straightforward, useful advice for clients. Hands-on experience running deep learning models on popular ML frameworks such as PyTorch, TensorFlow, ONNX Experience developing software solutions that run in Linux environments with containers and orchestration Experience with Source code and configuration management tools, Git knowledge is required. Customer-facing experience translating customer requirements into technical solutions (discovery, scoping, success criteria, and execution plans). Proven ability to build and deliver technical demos, proofs-of-concept, and reference applications for ML/GenAI workloads. Strong technical writing skills to produce customer-ready documentation (getting started guides, deployment runbooks, troubleshooting guides) and deliver partner training sessions. Experience driving issue triage and technical escalations with customers, coordinating across product, hardware, and software engineering teams to resolution. Excellent stakeholder management and communication skills: present complex technical concepts clearly to both engineering and non-engineering audiences.