Senior AI Engineer - Agentic AI & Knowledge Systems
About the Role
We're looking for a hands-on Senior AI Platform Engineer who has built and shipped production AI systems, not just prototypes. Working directly with the AI Engineering, Product, and Solutions teams, you'll design and build the core AI platform that powers agentic workflows, enterprise reasoning, retrieval systems, knowledge graphs, AI services, and intelligent automation. You'll own the architecture, engineering, and operational excellence of our AI platform, from model orchestration and inference to deployment, scalability, security, observability, and reliability.
Success in this role is measured by platform reliability, scalability, explainability, performance, cost efficiency, security, developer productivity, and customer impact.
What You'll Do
- Design, build, and operate production-grade AI platform capabilities from concept through enterprise-scale deployment.
- Build scalable multi-agent systems using frameworks such as LangGraph, CrewAI, AutoGen, or similar, including planning, reasoning, memory, tool use, orchestration, and state management.
- Design and implement enterprise retrieval systems using RAG, GraphRAG, semantic search, hybrid search, embeddings, reranking, vector databases, and knowledge graphs.
- Integrate and orchestrate multiple frontier, commercial and open-source foundation models, selecting the appropriate models based on accuracy, latency, cost, governance, and customer requirements.
- Build AI platform services for prompt management, model routing, tool orchestration, context management, session management, and AI workflow execution.
- Develop scalable inference APIs, AI microservices, SDKs, and backend services using Python, FastAPI, and modern cloud-native architectures.
- Design secure, multi-tenant AI systems with enterprise authentication, RBAC, PII protection, governance, compliance, auditability, and customer data isolation.
- Build AI evaluation infrastructure, guardrails, structured outputs, human-in-the-loop workflows, prompt versioning, observability, and production monitoring using tools such as LangSmith, Langfuse, OpenTelemetry, or similar.
- Optimize inference performance through model routing, caching, batching, quantization, parallel execution, GPU utilization, and cost optimization techniques.
- Build resilient AI systems capable of handling hallucinations, prompt regressions, model drift, service failures, fallback strategies, and operational recovery.
- Collaborate closely with Machine Learning Scientists to integrate trained models, evaluation frameworks, and continuous model improvement into production systems.
- Work with Knowledge Engineers to integrate semantic models, knowledge graphs, metadata, and enterprise context into AI workflows.
- Partner with Product and Solutions teams to translate customer requirements into scalable AI platform capabilities.
- Evaluate emerging AI technologies, frameworks, protocols, and architectures, making pragmatic engineering decisions that balance innovation with long-term maintainability.
- Contribute to the technical vision, engineering culture, coding standards, platform architecture, and long-term AI strategy of Ontio.
What We're Looking For
Required Qualifications
- 8+ years of experience building production AI applications, including 3+ years delivering Agentic AI or LLM-powered systems in production.
- Strong software engineering skills with Python and experience building scalable, maintainable production systems.
- Proven experience designing and building production AI applications, LLM-powered systems, AI agents, and enterprise AI platforms, not simply integrating AI APIs.
- Hands-on experience building autonomous or multi-agent systems using LangGraph, LangChain, CrewAI, AutoGen, Semantic Kernel, or similar frameworks.
- Experience with MCP, A2A, or emerging agent communication standards is highly desirable.
- Strong experience with Retrieval-Augmented Generation (RAG), GraphRAG, semantic search, embeddings, reranking, vector databases (Qdrant, Pinecone, Weaviate, Milvus, pgvector, or similar), and retrieval optimization.
- Experience integrating multiple LLM providers (OpenAI, Anthropic, Gemini, Azure OpenAI, open-source models, or similar) with production-grade model routing, evaluation, and fallback strategies.
- Experience building AI evaluation infrastructure, guardrails, prompt management, observability, structured outputs, and production monitoring.
- Strong backend engineering experience with FastAPI, REST APIs, asynchronous programming, distributed systems, Docker, Kubernetes, CI/CD, and cloud platforms such as AWS, Azure, or Google Cloud.
- Experience designing enterprise AI platforms with governance, security, RBAC, PII handling, compliance, auditability, and multi-tenant architectures.
- Strong understanding of system architecture, distributed systems, scalability, reliability, performance optimization, and cloud-native application design.
- Experience working with event-driven architectures, messaging systems, caching, and high-throughput APIs.
- Familiarity with model serving frameworks such as vLLM, Ollama, or similar technologies.
- Strong understanding of AI infrastructure, inference optimization, GPU utilization, model lifecycle management, and production AI operations.
- Excellent communication skills with the ability to work effectively across engineering, product, data science, and customer-facing teams.
- Ability to thrive in a fast-moving startup environment with strong ownership, autonomy, adaptability, and a bias toward execution.
Nice to Have
- Experience fine-tuning or serving LLMs using LoRA, QLoRA, PEFT, Hugging Face Transformers, or similar technologies.
- Experience with MLflow, Kubeflow, LangGraph Platform, LangSmith, Langfuse, OpenTelemetry, or similar AI platform tooling.
- Experience building enterprise SaaS platforms supporting BYOC, customer VPC, on-premises, or hybrid deployment models.
- Experience with Go, Rust, or high-performance backend development.
- Contributions to open-source AI projects, technical publications, patents, or developer communities.
How to Apply
Email your resume and a short note on why you're interested to careers@ontio.ai.
careers@ontio.ai