AI Platform Engineer – LLM and Agentic Systems

Job Category: Data AI & ML Jobs
Job Type: Remote

Job Description

A leading organization is seeking an AI Platform Engineer with expertise in large language model platforms, agentic workflows, and secure AI solution architecture. This position focuses on designing, customizing, and optimizing AI-powered development workflows while serving as a bridge between technical teams and business stakeholders.

The ideal candidate has hands-on experience building autonomous or multi-step AI systems, implementing secure model access patterns, and supporting enterprise-scale AI platforms using cloud-native technologies.

Responsibilities

  • Design and customize AI-assisted development environments and coding workflows.
  • Develop agentic AI solutions that incorporate iterative reasoning, automated error handling, validation processes, and code quality controls.
  • Build and maintain secure integrations with large language model platforms and cloud-based AI services.
  • Create workflows that incorporate regression testing, integration testing, and validation of AI-generated outputs.
  • Design and support proxy and gateway architectures that manage model access, security controls, and service routing.
  • Implement solutions that enforce API governance, credential protection, and operational guardrails.
  • Collaborate with engineering teams and business stakeholders to gather requirements and translate objectives into scalable AI solutions.
  • Evaluate and improve developer experience through enhanced tooling, automation, and workflow optimization.
  • Document solution architecture, operational procedures, and best practices for AI platform usage.

Required Experience and Skills

  • Deep expertise in Claude Code customization techniques and related tooling.
  • Advanced experience designing and implementing agentic prompting strategies and autonomous AI workflows.
  • Demonstrated experience building loop-based coding workflows, including automated error-correction processes and testing-driven validation for AI-generated code.
  • Strong understanding of AI security principles and secure platform design.
  • Experience with AWS Bedrock and AgentCore.
  • Experience building secure proxy architectures that manage API access, enforce guardrails, and support scalable model consumption.
  • Proven ability to work effectively with both technical and non-technical stakeholders.
  • Strong communication, collaboration, and stakeholder management skills.

Required Technical Validation Areas

  • Experience creating agentic workflows that extend beyond single-prompt interactions.
  • Experience implementing regression testing and automated validation within AI-assisted development processes.
  • Experience designing security-focused AI platform architectures.
  • Experience managing model access, credential handling, and endpoint routing within enterprise AI environments.

Preferred Experience and Skills

  • Experience with LiteLLM gateway implementations.
  • Experience supporting cloud-based coding agent workflows.
  • Strong background in regression testing and integration testing practices.
  • Experience improving developer productivity through developer experience initiatives.
  • Familiarity with scalable AI platform operations and governance frameworks.

FAQ

1. What does an AI Platform Engineer specializing in LLM and agentic systems do?

An AI Platform Engineer builds and operates the infrastructure, services, and tooling required to run large language model and agentic AI applications reliably in production. The role focuses on model integration, AI orchestration, scalable inference, platform automation, observability, security, and developer enablement.

2. What are agentic AI systems?

Agentic AI systems use models to perform multi-step tasks by planning actions, interacting with tools or external systems, and responding to changing inputs. Engineers build the platform capabilities that allow these agents to operate securely, reliably, and within defined business workflows.

3. How are LLMs integrated into enterprise applications?

LLMs can be connected to APIs, enterprise data sources, applications, tools, and workflow services to support tasks such as content generation, information retrieval, summarization, and decision support. The platform engineer helps provide the infrastructure and integration patterns needed to deploy these capabilities at scale.

4. What technologies are commonly used in this role?

The technology stack may include Python, cloud AI services, container platforms, Kubernetes, REST APIs, vector databases, model-serving systems, workflow orchestration tools, CI/CD platforms, and observability solutions. The specific stack depends on the organization’s AI architecture and cloud environment.

5. How does the role support AI model deployment and inference?

The engineer develops deployment workflows and runtime infrastructure for model inference, including scaling, resource management, API integration, monitoring, and version control. The goal is to provide reliable model access while managing performance, availability, and infrastructure costs.

6. What platform capabilities are important for agentic AI applications?

Agentic applications typically require secure tool integration, workflow orchestration, memory or state management, model routing, access controls, logging, evaluation, and failure handling. A well-designed platform provides reusable components so development teams can build agents without recreating core infrastructure.

 

Apply for this position

**If you have already submitted your resume for another Job Opening please do not re-apply to a different role. You can email through Contact Us about your interest in other roles.

Allowed Type(s): .pdf, .doc, .docx