LLM Inference & Efficiency Engineer100k以上
北京硕士及以上经验不限vLLMSGLang
带薪年假#五险一金#混合办公#节日福利 #零食下午茶
About the Group/Team
At Canva, we're building a future powered by AI that's as magical as it is impactful. As a Research Scientist for AI Efficiency at Canva, you'll be responsible for advancing the future of AI by experimenting with cutting-edge techniques, as well as improving models for real-world quality and performance.
About the Role
Canva is searching for outstanding AI Researchers to join the AI Efficiency team in Canva. We are dedicated to the development and optimization of GPU-accelerated efficient AI computing, for design generation, video generation, and more. We focus on pushing the boundary of creating AI models that are not only powerful, bot also computationally efficient. You will be part of this amazing team that tackles AI efficiency from low-level kernel optimization to high-level model distillations, from novel model architecture exploration to creating efficient inference systems. Your contributions will create real impact to Canva’s core products, used by hundreds of millions of active users.
At the moment, this role is focused on:
1. Designing and advancing state-of-the-art generative AI models across multimodal domains
2. Improving AI efficiency across the stack, including:
- Scalable inference systems
- Kernel and graph optimization
- Model compilation and systems optimization
- Model compression techniques (quantization, distillation)
- Efficient model architecture design
3. Collaborating cross-functionally with research, engineering, and product teams to bring innovations into production
4.Translating research breakthroughs into impactful features within Canva’s core products
5. Contributing to the broader AI community through publications at top-tier conferences
What we're looking for
You’re probably a match if you have:
1. You have strong expertise in foundation models, particularly in generative AI (e.g., diffusion models, LLMs, VLMs)
2. You bring experience in AI efficiency, such as:
- Efficient architecture design and inference systems
- Low-level optimization (kernel/graph)
- Model optimization (quantization, distillation, compression)
3. You have a strong academic and/or industry track record, including publications in leading conferences (e.g., NeurIPS, ICML, ICLR, CVPR) or meaningful open-source contributions
4. You’re proficient in tools such as Python, PyTorch, Transformers, Diffusers, Megatron, DeepSpeed, and cloud platforms
5. You’re a clear communicator who can collaborate effectively across teams