Job Title: Site Reliability Engineer (SRE)
Location: Remote
Duration: Long Term
Job Description
We are seeking a Site Reliability Engineer (SRE) with strong expertise in Splunk and production support to ensure the reliability, performance, and availability of cloud-based applications.
Required Skills
- 8+ years of experience as an Site Reliability Engineer
- Strong hands-on experience with Splunk (SPL, dashboards, alerts, log analysis, and monitoring).
- Experience with AWS and/or Azure cloud platforms.
- Hands-on experience with Docker and Kubernetes.
- Strong scripting skills using Python, Bash, or PowerShell.
- Experience with CI/CD tools such as Jenkins, Azure DevOps, or GitHub Actions.
- Solid understanding of SLIs, SLOs, Incident Management, and Root Cause Analysis (RCA).
- Strong Linux administration and production troubleshooting skills.
- Experience with ServiceNow integration is a plus.
Responsibilities
- Build and maintain Splunk dashboards, alerts, and monitoring solutions.
- Monitor production environments and resolve performance issues.
- Perform incident response, RCA, and automate operational tasks.
- Collaborate with development teams to improve system reliability and scalability.
Spruce Technology, Inc. is a mid-size, award-winning (Inc 5000, SmartCEO, Entrepreneur of the Year) technology services firm with a steadily growing portfolio of commercial and government clients. Spruce provides innovative technology solutions, specialized IT staff, and IT strategy consulting nationwide. Spruce maintains partnerships with major technology vendors and continually develops leading-edge offerings in service areas such as digital experience, data services, application development, infrastructure, cyber security, and IT staffing.
—