
C2C contract jobs
Location: Mountain View, CA (Onsite)
Contract
Key Responsibilities
• Build and extend O11Y pipelines: metrics, distributed tracing, logging, SLO/SLI definitions and dashboards consumed platform-wide.
• Instrument AI systems in production: token usage and cost, latency-per-inference, tool-call success rates, response-quality signals, agent loop detection, runaway-cost alerting.
• Extend tracing across agentic flows — planner → executor → retrieval → tool calls — spanning GenOS, AI Gateway, and MCP Gateway.
• Define SLOs where “available” includes response quality and tool-call success, not just HTTP 200s.
• On-call and incident command for AI platform surfaces: model/provider degradation, Bedrock/Gemini failover, semantic-cache issues, prompt-injection events — blameless postmortems with tracked remediations.
• Build AIOps agents in the Traffic Agent pattern: intelligent triage, anomaly detection, auto-remediation with hard guardrails.
• Drive UX Availability (R30) and Foundations Capabilities (R30); land network, O11Y, and cloud cost savings (tracked YTD KPIs).
Must-Have Qualifications
• 7+ years SRE/production engineering on Tier-0/Tier-1, high-traffic systems.
• Strong Go or Python — this is a build role: tooling, automation, instrumentation.
• Has built observability stacks, not just consumed them: Prometheus/Grafana, OpenTelemetry, or equivalent at scale, including cardinality and cost control.
• Production LLM/ML monitoring: Langfuse, Arize, WhyLabs, or homegrown — token/cost tracking, drift and quality metrics.
• Working fluency in AI-system failure modes: nondeterminism, provider limits and outages, context-window overflow, agent loops, cache poisoning.
• Kubernetes + AWS operational depth — debugs across cluster, mesh, and gateway layers.
• Structured incident-command and postmortem experience.
Nice-to-Have
• AIOps / LLM-applied-to-ops: auto-triage, incident summarization, remediation agents.
• Chaos engineering (Litmus, Gremlin, or homegrown); eBPF or deep network debugging.
• FinOps / cost engineering; fintech or regulated-industry reliability experience.
To apply for this job email your details to AjithG@Vbeyond.com