- Hybrid
- On-site
... preferred) - Expertise in LLM optimization: prompt engineering, prompt fine-tuning, embeddings, RAG, token management, and inference optimization - Track record of technical leadership: driving architecture decisions, mentoring engineers, and shipping complex projects - Experience with high-throughput inference using Ray, Triton, ...
Bengaluru, भारत