SRE with AI (Don’t share any AI/ML profiles) – C2C jobs

Contract

C2C contract jobs

SRE with AI

Location: Mountain View, CA (Onsite)

Contract

 

Key Responsibilities

•      Build and extend O11Y pipelines: metrics, distributed tracing, logging, SLO/SLI definitions and dashboards consumed platform-wide.

•      Instrument AI systems in production: token usage and cost, latency-per-inference, tool-call success rates, response-quality signals, agent loop detection, runaway-cost alerting.

•      Extend tracing across agentic flows — planner → executor → retrieval → tool calls — spanning GenOS, AI Gateway, and MCP Gateway.

•      Define SLOs where “available” includes response quality and tool-call success, not just HTTP 200s.

•      On-call and incident command for AI platform surfaces: model/provider degradation, Bedrock/Gemini failover, semantic-cache issues, prompt-injection events — blameless postmortems with tracked remediations.

•      Own the SRE side of the progressive-delivery seam: canary analysis, automated rollback decisioning, chaos/resilience testing, blast-radius controls — DevOps builds the pipeline; you define the gates.

•      Build AIOps agents in the Traffic Agent pattern: intelligent triage, anomaly detection, auto-remediation with hard guardrails.

•      Drive UX Availability (R30) and Foundations Capabilities (R30); land network, O11Y, and cloud cost savings (tracked YTD KPIs).

Must-Have Qualifications

•      7+ years SRE/production engineering on Tier-0/Tier-1, high-traffic systems.

•      Strong Go or Python — this is a build role: tooling, automation, instrumentation.

•      Has built observability stacks, not just consumed them: Prometheus/Grafana, OpenTelemetry, or equivalent at scale, including cardinality and cost control.

•      Production LLM/ML monitoring: Langfuse, Arize, WhyLabs, or homegrown — token/cost tracking, drift and quality metrics.

•      Working fluency in AI-system failure modes: nondeterminism, provider limits and outages, context-window overflow, agent loops, cache poisoning.

•      Kubernetes + AWS operational depth — debugs across cluster, mesh, and gateway layers.

•      Structured incident-command and postmortem experience.

Nice-to-Have

•      AIOps / LLM-applied-to-ops: auto-triage, incident summarization, remediation agents.

•      Chaos engineering (Litmus, Gremlin, or homegrown); eBPF or deep network debugging.

•      FinOps / cost engineering; fintech or regulated-industry reliability experience.

To apply for this job email your details to AjithG@Vbeyond.com

×

Post your C2C job instantly

Quick & easy posting in 10 seconds

Keep it concise - you can add details later
Please use your company/professional email address
Simple math question to prevent spam