
C2C contract jobs
Bellevue WA
Contract
About the Role
We are looking for a Technical Project Manager (TPM) to lead the Tier 3 Production Support function for the T-Life BFF (Backend for Frontend) platform at Client. This is a high-visibility, high-impact role at the intersection of production operations, multi-vendor coordination, and engineering excellence.
The T-Life BFF platform powers Client’s flagship customer-facing digital experiences. As the Tier 3 TPM, you will own the escalation path for the most complex production issues, coordinate across multiple vendor teams managing 5 development pods, and drive resolution while maintaining 100% platform availability. You will work in an environment where production incidents and new feature delivery happen simultaneously, timelines are aggressive, and the team operates 24/7.
This is not a traditional PM role. You must be technically hands-on — comfortable triaging API issues in Splunk, understanding distributed system failures, and driving root cause analysis alongside engineers.
Key Responsibilities
Tier 3 Production Support Ownership
• Own the Tier 3 escalation process for T-Life BFF platform — serve as the final point of resolution for complex production issues
• Lead war rooms and incident bridges for P1/P2 incidents, driving cross-team coordination until resolution
• Perform hands-on API triage using Splunk — analyze logs, trace request flows, identify root cause across BFF and downstream microservices
• Drive blameless post-incident reviews (PIRs) and ensure corrective actions are tracked to closure
• Maintain and enforce SLA compliance for incident response and resolution times
• Build and maintain Splunk dashboards for real-time API health monitoring, error rate tracking, and latency analysis
Multi-Vendor Program Coordination
• Coordinate production support activities across 5 development pods managed by external vendor(s)
• Establish and run cross-vendor governance cadences — daily standups, weekly syncs, escalation protocols
• Drive accountability for production fix timelines when dev resources are not under direct management
• Manage dependencies and handoffs between platform engineering (UST) and development vendor teams
• Facilitate conflict resolution between vendors during production incidents without finger-pointing
• Track and report vendor performance metrics — MTTR, incident recurrence, fix quality
Stakeholder & Product Management
• Partner with Product team to balance production stability needs against feature delivery timelines
• Negotiate realistic timelines when product requirements conflict with available engineering capacity
• Provide transparent, data-backed status updates to T-Life leadership on production health and risks
• Implement capacity allocation models (production support vs. feature work) and make trade-offs visible to stakeholders
• Shield engineering teams from scope creep and unrealistic mid-sprint additions through structured change control
Operational Excellence & Continuous Improvement
• Identify recurring production issues and drive systemic fixes — break the cycle of repeated hotfixes
• Champion automation initiatives to reduce manual toil in incident response and triage
• Establish and track DORA metrics, incident trends, and team health indicators
• Implement on-call rotation and workload management practices to prevent team burnout during 24/7 operations
• Drive adoption of best practices — canary deployments, feature flags, pre-prod validation gates — to reduce deployment-related incidents
AI-Driven Engineering & Agentic Solutions
• Leverage AI coding assistants (Claude Code, Claude API) to accelerate incident triage, automate log analysis, and generate root cause hypotheses from Splunk data
• Design and implement agentic AI workflows — autonomous multi-step solutions that integrate with Jira, Confluence, Splunk, and CI/CD pipelines to automate repetitive engineering and operational tasks
• Build AI-powered automation for production support — intelligent alert correlation, automated runbook execution, and predictive incident detection using LLM-based agents
• Drive adoption of prompt engineering best practices and MCP (Model Context Protocol) server integrations to connect AI agents with enterprise tools and data sources
• Evaluate and implement AI-assisted code review, documentation generation, and test automation to improve engineering velocity across vendor teams
• Champion responsible AI adoption — establish guardrails, review AI-generated outputs for accuracy, and ensure compliance with client security and data governance policies
Required Qualifications
Experience
• 8+ years of experience in technical project/program management in software engineering environments
• 3+ years managing production support / Tier 3 operations for large-scale, customer-facing platforms
• 3+ years of multi-vendor program management experience — coordinating deliverables across 2+ vendor organizations
• Proven track record of managing production incidents (P1/P2) in 24/7 environments with aggressive SLAs
• Experience working in telecom, digital commerce, or large-scale consumer platform environments preferred
Technical Skills
Area Required Proficiency
Observability & Monitoring Splunk (must-have), SignalFx, Prometheus, Grafana Hands-on
API & Microservices RESTful APIs, BFF architecture, distributed tracing, API gateway patterns Working knowledge
Cloud & Infrastructure AWS (EKS, EC2, S3, CloudWatch), Kubernetes, Helm Working knowledge
CI/CD & DevOps GitLab CI/CD, deployment pipelines, canary/blue-green deployments Conceptual +
Incident Management PagerDuty / ServiceNow / Jira Service Mgmt, war room facilitation Hands-on
Project Management Jira, Confluence, Agile/Scrum, SAFe (preferred) Expert
AI & Agentic Solutions Claude Code / Claude API, Agentic AI workflows, LLM-powered automation, prompt engineering, MCP servers Hands-on
Core Competencies
• Incident Triage & RCA: Ability to read Splunk logs, trace API call chains, differentiate BFF-layer vs. downstream failures, and correlate deployment events with production regressions
• Multi-Vendor Negotiation: Influence without authority — drive accountability with vendor dev teams where you don’t have management control
• Stakeholder Communication: Translate technical issues into business impact for leadership; provide options-based recommendations, not just problem statements
• Prioritization Under Pressure: Make real-time trade-off decisions during simultaneous production incidents and feature deadlines
• Team Sustainability: Proactively manage workload, rotate on-call, and prevent burnout in a 24/7 operating environment
• Data-Driven Decision Making: Use velocity data, incident trends, MTTR metrics, and capacity models to negotiate timelines and resource allocation
• AI & Automation Mindset: Hands-on experience with AI coding tools (Claude Code/API), building agentic solutions that autonomously execute multi-step workflows, and driving AI-first approaches to reduce manual toil in production operations
Preferred Qualifications
• Experience with client technology ecosystem or large telecom platforms
• Familiarity with BFF (Backend for Frontend) architectural patterns and API gateway technologies
• Experience with AI/ML-powered operations tools (AIOps, intelligent alerting, automated RCA) and hands-on expertise building agentic AI solutions using Claude Code, Claude API, or similar LLM platforms for engineering automation
• PMP, CSM, SAFe Agilist, or ITIL certification
• Experience building Splunk dashboards and automated alerting rules
• Experience designing and deploying agentic AI workflows that integrate with enterprise tools (Jira, Confluence, Splunk, CI/CD) via MCP servers or API integrations
• Demonstrated ability to leverage AI assistants for code generation, automated documentation, test creation, and operational runbook automation
To apply for this job email your details to Tejaswini.b@metasisinfo.com