Loading open roles
Loading open roles
Loading role
Recroots · posted 5 months ago
Architect enterprise-level systems emphasizing reliability, scalability, availability, and operational excellence (primarily AWS). • Define, build, and institutionalize SRE frameworks for SLIs/SLOs, error budgets, reliability scorecards, and AWS Operational Readiness Reviews to meet or exceed reliability, security, and cost-efficiency targets. • Set technical direction for CI/CD, containerization (Docker, Kubernetes), and IaC (Terraform, CloudFormation); establish organization-wide standards and ensure adoption and compliance across product teams. • Architect advanced monitoring, observability, and incident management; lead RCAs, blameless postmortems, and continuous improvement using Datadog, Prometheus, Grafana, and ELK. __________________________________________________________________________________________ • Design and roll out AI-powered platforms and processes (e.g., Copilot, ChatGPT, custom AI frameworks) to improve developer productivity, incident response, and operational intelligence. • Provide deep technical consultation and embedded enablement for product teams, including architectural reviews, operational readiness, reliability training, and documentation (flow/class/sequence/data-flow diagrams, database schemas, component docs, wireframes). • Create reusable frameworks through hands-on development and testing; guide technology direction, recommend platforms, and steward the Cvent Paved Road with aligned training and staff development. • Partner with security, IT, and business operations; drive stakeholder alignment, manage technical debt and refactoring, and build support for proposals and solutions. • Serve as the primary evangelist and mentor for SRE best practices, reliability culture, and technical excellence across divisions. • Help guide the technology directions by recommending specific technologies to pursue, suggesting training and staff development activities, and monitoring the Cvent Paved Road • Keep an eye on any increase technical debt and propose remedies and code refactoring • Work closely with the implementation teams of the various product lines to ensure that architectural standards and best practices are being followed consistently across all Cvent applications • Stakeholder management with an ability to gain others’ support for ideas, proposals, projects, and solutions What You Will Need for this Position: • Bachelor’s degree required; Master’s preferred in Computer Science, Computer Engineering, or related field. • 14+ years in SRE, cloud architecture, or infrastructure roles with significant leadership and organizational impact. • Deep hands‑on expertise in AWS (multi‑account, multi‑region, advanced networking, security, cost optimization); AWS Solutions Architect Professional/Advanced certifications. Multi‑cloud/migration experience a plus. • Mastery of CI/CD, DevOps toolchains, containerization, and infrastructure‑as‑code; proven track record implementing SRE frameworks (ORRs, SLIs/SLOs, reliability reviews) and driving cross‑team maturity. • Exceptional programming in Python/Go (or similar) and systems design; strong troubleshooting and scaling of distributed systems. • Strong knowledge and practical implementation of AI/automation for operational efficiency and developer productivity. • Deep experience in security and compliance for regulated industries; ability to design secure solutions aligned to industry best practices. • Proficiency with relational and NoSQL databases (e.g., PostgreSQL, MySQL, SQL Server, Oracle, Redis, Couchbase, Elasticsearch). • Excellent communication and influence across executives, architects, engineers, and external stakeholders; ability to articulate complex systems clearly. • Ability to estimate work, partner with Project Managers on task‑level plans/proposals, and deliver predictably. • Self‑motivated, operates with minimal supervision; publishes/speaks or contributes to open source in SRE/DevOps/reliability.