# Elice Infrastructure SRE Engineer - Remember 331349
AI Summary
Purpose:
- Preserve the Elice Infrastructure SRE Engineer posting used for the 2026-08-05 tailored application package.
Key points:
- The role covers SLI/SLO-based reliability operations, Prometheus/Grafana/ELK observability, incident response, RCA, architecture and deployment improvements, and toil removal with scripts or IaC.
- Required qualifications are observability-stack experience or equivalent knowledge, Python or Shell automation, and calm incident analysis and response.
- Preferred qualifications add Linux/Kubernetes/TCP-IP operations, GitOps or CI/CD, Chaos Engineering, cloud infrastructure, SLI/SLO and Error Budget operations, and FinOps/AIOps project leadership.
- The candidate directly matches observability, automation, Kubernetes/Linux operations, and incident RCA. Formal SLI/SLO, Error Budget, Chaos Engineering, production GitOps/IaC, FinOps/AIOps, and GPU/InfiniBand/Ceph operations remain gaps.
Relevant when:
- Reviewing or revising the Elice application under
career/맞춤이력서/엘리스-SRE-엔지니어/. - Preparing for the phone screen, mini project, or interviews about SRE concepts and existing experience boundaries.
Do not read full document unless:
- Exact posting wording, process, benefits, or infrastructure-team context is needed.
Linked documents:
../../wiki/projects/data-platform-systems-engineering.md../../wiki/projects/infra-db-monitoring.md
Open Questions
- Exact closing date and salary band: Needs confirmation.
- On-call rotation, call frequency, compensation, and compensatory time: Needs confirmation.
- Remote and flexible-work policy: Needs confirmation.
- Current SLO and Error Budget adoption level and SRE team size: Needs confirmation.
- Expected mini-project duration: Needs confirmation.
Details
Source: <https://career.rememberapp.co.kr/job/posting/331349>
Captured: 2026-08-05 via Safari-compatible HTTP request and visible-text extraction.
Career: 3-10 years
Education: bachelor's degree or higher
Location: Seoul, Gangnam-gu
Salary: negotiable
Deadline: closes when hired
Remember context at capture: Series C, accumulated investment above KRW 33.3B, 51-300 employees, KRW 500K hiring reward.
Company and team context from the posting
- Elice presents itself as an AI company covering infrastructure, platform, models, and content.
- The infrastructure team description includes AI PMDC, ECI, GPU servers, Ceph, parallel filesystems, Kubernetes, networking, security, and data-center layers.
- Water-cooled B200 GPUs, InfiniBand 400G, 100G-class server/storage/ISP networks, and hundreds-node GPU clusters are team context, not candidate experience.
Responsibilities
- Define latency, error-rate, and other indicators and design operating processes for SLI/SLO-based reliability.
- Build, operate, and optimize Prometheus, Grafana, ELK, and related monitoring, alerting, and logging platforms.
- Lead incident-response processes and reduce MTTR.
- Perform RCA and fix structural problems in system architecture and deployment pipelines.
- Identify toil and remove it with scripts and IaC.
Required qualifications
- Experience building or operating monitoring, alerting, logging, or an equivalent observability stack, or equivalent knowledge for junior candidates.
- System operations and automation using Python, Shell, or equivalent scripting languages.
- Experience calmly analyzing and responding to service incidents.
Preferred qualifications
- Deep Linux, Kubernetes, and TCP/IP knowledge with at least three years of operations experience.
- Prometheus, Grafana, ELK/Loki, or comparable open-source observability operations.
- Deployment-stability improvements using GitOps or CI/CD pipelines.
- Chaos Engineering adoption or execution.
- AWS, Azure, or GCP infrastructure operations.
- SLI/SLO design and Error Budget-based operations.
- FinOps or AIOps infrastructure-project leadership.
Process
Document review -> phone screen -> mini project -> in-person interview -> reference check -> compensation negotiation -> final acceptance.
Benefits listed in the posting
- Current work devices, motion desks, and welcome kit.
- Claude Code, e-library, Elice courses, and job-related education.
- Company cafeteria lunch, dinner support, late-night taxi, and snacks.
- All-hands meetings, M365/Teams/Outlook, team scrums, casual dress, and
님titles. - Performance incentives, monthly recognition, and team dinners.
- Family-event support and leave, vaccination and health-screening leave.
- Mingle lunch, buddy program, and leader communication budget.