CURRENT OPENINGS
Lead Platform Engineer
We welcome individual candidates and Corp-to-Corp (CTC) resume submissions.
6-month contract
NO DIRECT CALLS OR EMAILS! WILL NOT BE RETURNED
We are seeking a Lead Platform Engineer to help our direct client. This is a hybrid position. Must live within 50 miles of Baltimore, MD, Wilmington, DE, Charlotte, NC, Dallas, TX, New York, NY, or Evansville, IN.
The Lead Platform Engineer will be part of a high‑performing Monitoring Engineering team within a fast‑paced financial technology organization. In this role, this person will apply SRE principles to design, build, and evolve monitoring and observability capabilities that ensure the reliability, performance, and operability of core applications and infrastructure.
They will partner closely with application, platform, and development teams to implement data‑driven alerting, SLO/SLA-based monitoring, telemetry pipelines, dashboards, correlations, and automated remediation. Will work directly to improve system reliability, reduce MTTR, and enhance enterprise‑wide operational insight. This role requires strong analytical thinking, systems engineering discipline, and a proactive approach to identifying risks, preventing incidents, and driving continuous improvement across the production ecosystem.
Responsibilities
- Architect, deploy, and operate OpenTelemetry‑based telemetry pipelines, including instrumentation standards, collector configurations, sampling strategies, and routing to Elastic and other backends.
- Develop and maintain instrumentation, telemetry, and alerting for the Enterprise Monitoring Center using industry‑leading tools, such as:
- Grafana, OpsRamp, ElasticStack, BigPanda
- AWS CloudWatch, Azure Monitor
- Drive observability standards and best practices across multiple engineering teams through influence, documentation, and partnership rather than direct authority.
- Apply SRE best practices to ensure measurable SLIs/SLOs, reliability dashboards, and health indicators for critical systems.
- Integrate and manage OpenTelemetry for distributed tracing and telemetry data collection, enabling end‑to‑end visibility of business‑critical transactions.
- Collaborate with application development teams to define and document observability requirements for each project or release, ensuring accurate and actionable monitoring and tracing are in place for every step of business‑critical workflows.
- Embed reliability considerations early in the SDLC, including SLO definitions, instrumentation needs, and failure‑mode awareness.
- Partner with product and engineering teams to use SLOs and error budgets to guide release decisions, prioritization, and toil reduction.
- Define and maintain standardized alert payloads per engineering guidelines, ensuring alerts are actionable.
- Partner with Level 2 and Level 3 support teams to reflect process changes in monitoring dashboards.
- Maintain and optimize thresholds, ensuring seamless escalations via BigPanda as the central alert hub.
- Dashboard Creation & Maintenance
- Develop and maintain technical documentation, runbooks, diagnostic guides, and observability standards across the enterprise.
- Evaluate and refine release, deployment, and monitoring processes to support consistent, reliable delivery pipelines.
- Mentor junior engineers and promote a culture focused on reliability, automation, and operational excellence.
- Build automation frameworks for monitoring, alerting, self‑healing workflows, and incident response to reduce toil and improve MTTR.
- Drive system optimization through capacity analysis, performance tuning, and proactive detection of reliability risks.
- Contribute to the automation of routine operational tasks to improve system reliability and engineer quality of life.
- Advocate for and implement observability best practices across engineering teams.
- Define, implement, and operationalize SLIs, SLOs, and error budgets for critical services.
- Participate in and improve incident response processes, including detection, triage, escalation, and recovery.
- Bachelor’s degree in computer science, IT, or related field.
- 5+ years of experience in software, systems, or reliability engineering roles, with multiple years of hands‑on experience owning production observability, monitoring, and SLOs in distributed systems.
- Expertise building scalable, reliable monitoring and observability solutions, including instrumentation, alerting, dashboarding, and configuration across large, complex environments.
- Hands‑on expertise and proficiency with modern monitoring and observability tools, (e.g., OpsRamp, Grafana, Elastic, CloudWatch, Azure Monitor BigPanda (AIOps), and strong knowledge of metrics, logs, traces, and OpenTelemetry.
- Strong scripting and programming capability (Bash, PowerShell, and one or more languages such as Python, C-family, or JavaScript) to automate telemetry, alerting, and platform workflows.
- Exceptional expertise with cloud platforms (AWS and/or Azure) and container orchestration systems (Kubernetes, Docker).
- Heavy hands‑on experience with Elastic Observability (APM, Logs, Metrics, Traces)
- Understanding of distributed systems fundamentals, including networking, security, databases, DevSecOps principles, and performance/capacity engineering.
- Ability to clearly explain complex technical topics to both technical and non‑technical audiences.
- Problem‑solving and troubleshooting abilities, especially in high‑pressure or time‑sensitive environments.
- Proven cross‑functional collaboration, working seamlessly with diverse teams in large, complex IT environments and driving continuous improvement across systems.
Nice to Have:
- Experience with CI/CD pipelines and tools like Jenkins, GitHub, GitLab CI, or CircleCI
- Experience querying, manipulating, and visualizing time‑series data.
- Familiarity with Infrastructure as Code tools (e.g., Ansible, Terraform).
- Knowledge of microservices architecture and event-driven systems.
- Working knowledge of REST APIs, JSON, and ServiceNow.
- Experience with cloud monitoring—particularly AWS or Azure.
We look forward to receiving your resume in PDF format.