company logo

Mercari

Software Engineer (Site Reliability)

゜フトりェア゚ンゞニアサむト信頌性

Tags: Full-time, 6~8 YOE, Business Japanese

Minato City, Tokyo, Japan・Fetched 30+ days ago

Job Description

Employment Type: Full-time
Team: Engineering

JD in Japanese follows. 英文の埌に和文JDをご芧いただけたす。


Software Engineer (Site reliability) - Mercari 

  • Employment Status: Full-time
  • Work Hours: Full Flextime (no core time) 
  • Office: Roppongi

For more details, see the Overview of Our Positions section on our Careers site.  



About Mercari

Circulate all forms of value to unleash the potential in all people

"What can I do to help society thrive with the finite resources we have?" The Mercari marketplace app was born in 2013 out of this thought by our founder Shintaro Yamada as he traveled the world. We believe that by circulating all forms of value, not just physical things and money, we can create opportunities for anyone to realize their dreams and contribute to society and the people around them. Mercari aims to use technology to connect people all over the world and create a world where anyone can unleash their potential. For more information about Mercari Group’s mission, see Mercari’s Culture Doc


 
Organization/Team Mission

Mercari Engineering Principles  

Mercari Engineering Principles are a shared understanding that serves as the foundation of engineering beliefs and behavior at Mercari. The Engineering Principles are designed to complement the organizational identity (Mercari’s mission, values, and culture) from an engineering viewpoint. 

These principles ultimately help us achieve Mercari’s mission by defining the ideal state we seek to realize in the long term. 

  • Passion For The Product
  • Grow Together
  • Solve Through Mechanisms
  • Collaborate Openly

For more details, please see the following link:

In the SRE Team drives the reliability, scalability, and operational excellence of Mercari Group’s production services, including Mercari, Merpay, and Mercoin. We make customer experience observable through CUJ SLOs and turn those insights into meaningful service improvements. Working across Google Cloud and Kubernetes, we strengthen incident response, reduce toil, and build resilient systems. As AI accelerates development, we are advancing guardrails, autonomous operations, and cross-team collaboration to prevent incidents, recover faster, and improve products end to end.

See here for more information about our mission and values.



Work Responsibilities

  • Operate hundreds of production microservices on Google Cloud (Kubernetes, managed services) under SLO targets, including on-support rotation for urgent issues.
  • Lead end-to-end reliability epics independently — from design through rollout, monitoring, and post-launch iteration.
  • Define and operate SLOs and SLIs for critical user journeys, using error budgets to prioritize work with product teams.
  • Lead incident response, advance the team's postmortem culture, and drive follow-ups that prevent recurrence.
  • Build autonomous AI agents for detection, triage, and recovery (log analysis, alert summarization, RCA, remediation) with clear safety.
  • Write Infrastructure as Code with Terraform and build automation to reduce toil and scale operations across a large microservices environment.
  • Build and maintain monitoring, alerting, and tracing on Datadog tied to user impact, with short detection-to-mitigation time.
  • Perform reliability and performance tuning on production workloads (capacity planning, autoscaling, load shedding, dependency hardening).
  • Partner with product and platform teams on production readiness reviews, capacity planning, and new infrastructure adoption.
  • Strengthen reliability governance through engineering (risk assessment, audit response, compliance-as-code).


 
Unique Challenges

  • Driving reliability across a wide business portfolio (Marketplace and Fintech) at Mercari scale, through a CUJ SLO that grounds product decisions and operational policy in user impact.
  • Working at the frontline of AI-driven development velocity and the operational strengthening it demands — you will help shape the SRE culture that emerges as engineers and autonomous agents share the work.
  • Partnering with engineering teams that own reliability and resilience and act on data, so reliability improvements compound across the company.
  • Balancing ~50% reactive work (alerts, on-support tickets, operations, incident response) with project delivery.
  • Working in a bilingual environment: you use both Japanese and English every day, within the team and across teams.


 
Qualifications

  • Required Experience/Skills
    • Production SRE experience with service ownership, availability targets, toil reduction, and operational readiness, including using SLOs and SLIs to prioritize reliability work alongside development teams.
    • Experience operating production services at scale (>10K QPS, or owning multiple production microservices) under SLOs.
    • Production experience with Google Cloud (compute, networking, managed services) and Kubernetes-based workloads.
    • Infrastructure-as-Code experience (Terraform) and scripting in Go, Python, or shell.
    • Hands-on experience with monitoring and observability (Datadog or equivalent), including alert design and reducing alert fatigue.
    • Experience owning incident response, postmortems, and on-call or on-support rotations.
    • Ability to lead epics end-to-end without teammate support.
    • Willingness to learn and apply AI to operational workflows beyond your core SRE expertise.
  • Preferred Experience/Skills
    • Experience designing or running platform-wide SLO programs across multiple services or business units.
    • Experience applying AI to operational workflows (log analysis, alert summarization, runbook assistance, RCA, remediation) with evaluation of accuracy and safety.
    • Experience operating high-scale Kubernetes platforms, or distributed systems internals (scheduling, consistency, failure recovery).
    • Experience leading reliability or platform initiatives that span multiple teams.
    • Experience strengthening reliability governance through engineering (compliance-as-code, automated audit evidence, risk assessment).
  • Language 
    • Japanese: Independent (CEFR – C1)
    • OR English: Independent (CEFR – C1)

For details about CEFR, see here.


 

Learn More About Mercari Group



Recruiting at Mercari

At Mercari Group, we value empathizing with and embodying the mission and values ​​of the Group and each company. To promote the creation of an organization that maximizes the total amount of value exhibited by all members, we would like to understand the experience and skills of each candidate as accurately as possible.

Recruiting cycle at Mercari Group

  • Application screening
  • Skill assessment: For engineering positions, you will be asked to complete a skill assessment on HackerRank or GitHub. For non-engineering positions, you may be asked to complete an assessment depending on the position. (The timing of the assessment may coincide with the interview process.)
  • Interview: The number of interviews may vary depending on the position.
  • Reference check: We will ask for online references around the timing of the final interview.
  • Offer: Offers will be determined carefully in consideration of the final interview and the reference check.

 Learn more about our recruiting process here.


 
Equal Opportunity Hiring

Here at Mercari, we work to realize a world in which no one’s potential is limited by their background and everyone has the opportunity to freely create value. We also firmly believe that a mindset of Inclusion & Diversity is essential for us to achieve our mission.

This, of course, extends to our hiring practices as well. Mercari is committed to eliminating discrimination based on age, gender, sexual orientation, race, religion, physical disability, and other such factors so that anyone who shares our mission and values can join us, regardless of their background. For more details, please read our I&D statement.

Please read and acknowledge our Privacy Policy prior to submitting your application.





Software Engineer (Site reliability) - Mercari 

  • 雇甚圢態  正瀟員
  • 働き方 フレックスタむム制コアタむムなし・フレキシブルタむムなし
  • 勀務地 六本朚

詳现はキャリアサむトの募集芁項よりご確認ください


 
メルカリグルヌプに぀いお

あらゆる䟡倀を埪環させ、あらゆる人の可胜性を広げる

「地球資源が限られおいるなか、より豊かな瀟䌚を぀くるために䜕ができるか」。2013幎、創業者の山田進倪郎が䞖界䞀呚の旅で抱いた課題意識から、フリマアプリ「メルカリ」は生たれたした。私たちは、物理的なモノやお金に限らずあらゆる䟡倀を埪環させるこずで、誰もがやりたいこずを実珟し、人や瀟䌚に貢献するための遞択肢を増やすこずができるず信じおいたす。

テクノロゞヌの力で䞖界䞭の人々を぀なぎ、あらゆる人の可胜性が発揮される䞖界を実珟しおいきたす。メルカリグルヌプの目指すべき方針に぀いおは Mercari Culture Doc をご芧ください。

 

 
組織・チヌムのミッション

  •  Mercari Engineering Principles
    Mercari Engineering Principles は、メルカリの゚ンゞニアリング組織における信念や行動の基盀ずなる共通認識を明文化したもので、メルカリのメンバヌ党員が共有するMission、Value、Cultureを゚ンゞニアリングの芖点から補完するものずなりたす。これらのPrinciplesは、私たちが長期的に実珟しようずする理想的な姿を定矩するこずで、最終的にメルカリのミッションを達成するために掻甚しおいきたす。
  • Passion For The Product
  • Grow Together
  • Solve Through Mechanisms
  • Collaborate Openly

詳现に぀いおぱンゞニアリングカルチャヌ  をご芧ください
 

SREチヌムは、Mercari、Merpay、Mercoinを含むMercari Groupの本番サヌビスにおいお、信頌性、スケヌラビリティ、運甚品質の向䞊をリヌドするチヌムです。システムやサヌビス単䜍だけでなく、お客さたの䜓隓を起点ずした信頌性を重芖し、CUJ SLOを掻甚しお重芁な䜓隓を芳枬可胜にしながら、継続的なサヌビス改善に぀なげおいたす。Google CloudやKubernetesをはじめずするクラりドネむティブな基盀䞊で運甚される倧芏暡なシステムに向き合い、むンシデント察応の高床化、トむル削枛、レゞリ゚ントなシステム蚭蚈・運甚を掚進しおいたす。さらに、AIによっお開発スピヌドが加速する䞭で、ガヌドレヌルの敎備、AI Agentを掻甚した調査・察応プロセスの高床化、チヌム暪断のコラボレヌションを通じお、むンシデントの未然防止ず迅速な埩旧を支えおいたす。

  • メルカリのミッション・バリュヌに぀いおの詳现はこちらをご芧ください


 
業務内容

  • SLOに基づき、Google Cloud䞊で皌働する数癟芏暡の本番マむクロサヌビスを運甚し、緊急時のオンコヌル察応や運甚サポヌトも担う。
  • 信頌性向䞊に向けた取り組みを、蚭蚈、監芖、リリヌス、リリヌス埌の改善たで䞀貫しおリヌドする。
  • 重芁なナヌザヌゞャヌニヌに察するSLI/SLOを定矩・運甚し、゚ラヌバゞェットを掻甚しおプロダクトチヌムずの優先順䜍付けを行う。
  • むンシデント察応をリヌドし、ポストモヌテム文化を発展させ、再発防止に぀ながるフォロヌアップを掚進する。
  • ログ分析、アラヌト芁玄、根本原因分析、埩旧察応をなどを察象に、明確な安党性を担保したAI Agentを構築する。
  • TerraformによるInfrastructure as Codeを実践し、倧芏暡なマむクロサヌビス環境における運甚のスケヌラビリティ向䞊やトむル削枛のための自動化を実斜する。
  • お客さたの圱響に玐づいた監芖、アラヌト、オブザヌバビリティを構築・維持し、怜知から緩和たでの時間を短瞮する。
  • 本番環境で皌働するサヌビスに察しお、利甚増加や障害発生を芋据えたリ゜ヌス蚭蚈、オヌトスケヌル、䟝存先サヌビスぞのレゞリ゚ンス匷化などを通じお、信頌性ず性胜を継続的に改善する。
  • プロダクトチヌムおよびプラットフォヌムチヌムず連携し、新機胜や新芏サヌビスを安党に本番投入するための準備、運甚蚭蚈、新しいプラットフォヌム敎備を掚進する。
  • リスク評䟡、監査察応、運甚ルヌルのコヌド化・自動怜蚌を通じお、安党で信頌性の高い本番環境を継続的に維持・改善する。


 
ナニヌクなチャレンゞ

  • MarketplaceずFintechを含むMercari Groupの幅広い事業領域においお、CUJ SLOを通じおお客さたぞの圱響を可芖化し、プロダクト刀断や運甚方針に反映しながら、信頌性向䞊に取り組みたす。
  • AIによっお開発スピヌドが加速する䞭で、゚ンゞニアずAI Agentが協働する新しい運甚のあり方を぀くり、これからのSRE文化を圢づくりたす。
  • 信頌性ずレゞリ゚ンスの重芁性が組織党䜓で理解されおいる環境で、プロダクトチヌムやプラットフォヌムチヌムず連携し、デヌタに基づく改善を継続的に進められたす。
  • 運甚サポヌト、問い合わせ察応、むンシデント察応などの日々の運甚業務ず、信頌性向䞊に向けた䞭長期のプロゞェクト掚進の䞡方に取り組みたす。
  • チヌム内倖で日本語ず英語を日垞的に䜿う環境で働きたす。


 
応募芁件

  • 求める経隓・スキル
    • サヌビスの信頌性に責任を持ち、可甚性目暙の達成、トむル削枛、本番皌働に向けた準備を掚進した経隓。SLI/SLOを掻甚し、開発チヌムず連携しながら信頌性向䞊の優先順䜍を刀断した経隓を含む。
    • SLOに基づき、倧芏暡なサヌビス10K QPS以䞊、たたは耇数の本番マむクロサヌビスを運甚した経隓。
    • Google CloudなどのクラりドサヌビスおよびKubernetes䞊で皌働するワヌクロヌドの本番運甚経隓。
    • Infrastructure as Codeの実践やSRE業務向けツヌルの開発を通じお、運甚の効率化・自動化を掚進した経隓。
    • Datadogたたは同等のツヌルを甚いた監芖・オブザヌバビリティ匷化の実務経隓。アラヌト蚭蚈や疲劎の軜枛に取り組んだ経隓を含む。
    • むンシデント察応、ポストモヌテム、オンコヌルたたは運甚サポヌトの圓番制を担った経隓。
    • 信頌性向䞊に向けた取り組みを、蚭蚈から実行、改善たで自埋的にリヌドできる。
    • SREの専門領域に閉じず、AIを運甚業務に孊習・適甚しおいく意欲。
  • 歓迎する経隓・スキル
    • 耇数のサヌビスたたは事業領域をたたぐ、党瀟的・暪断的なSLOプログラムの蚭蚈たたは運甚経隓。
    • ログ分析、アラヌト芁玄、根本原因分析、埩旧察応などの運甚業務にAIを掻甚した経隓、およびその粟床や安党性を評䟡した経隓。
    • 倧芏暡なKubernetes基盀の運甚経隓、たたは分散システムの内郚動䜜に関する知識・経隓。
    • 耇数チヌムにたたがる信頌性向䞊たたはプラットフォヌム改善の取り組みをリヌドした経隓。
    • リスク評䟡、監査蚌跡の自動収集、むンフラ蚭定や運甚ルヌルのコヌド化・自動怜蚌を通じお、本番環境の安党性ず信頌性を高めた経隓。
  • 語孊力
    • 日本語Independent (CEFR – C1)
    • OR 英語Independent (CEFR – C1)

※CEFRの詳现に぀いおは、こちらをご芧ください

 

メルカリグルヌプに぀いお知る 


 
遞考に぀いお

メルカリグルヌプではメルカリグルヌプおよび各カンパニヌのミッションずバリュヌぞの共感・䜓珟を倧切にしおいたす。メンバヌが発揮する䟡倀の総量が最倧化されるような組織づくりを掚進するために、候補者のみなさんの経隓やスキルをより正しく理解したいず考えおいたす。

遞考の流れ

  • 曞類遞考
  • 技術課題゚ンゞニアポゞションではHackerRankたたはGithubでの技術課題を、゚ンゞニア以倖のポゞションでは採甚ポゞションによりたす面接タむミングず前埌するこずがありたす
  • 面接ポゞションにより、耇数回の面接をお願いしたす
  • リファレンスオンラむン回答圢匏のもので、最終遞考の前埌でお願いしたす
  • オファヌ最終遞考ずリファレンスの内容より決定されたす

 

 ※詳しくは  こちらのペヌゞをご芧ください

 

遞考における機䌚の平等  

メルカリでは、バックグラりンドによっお個人の可胜性が決め぀けられるこずなく、自由に䟡倀を生みだす機䌚を手にできる瀟䌚の実珟を目指しおいたす。そしおメルカリがミッションを実珟するために「Inclusion & Diversity」ずいう考え方は䞍可欠な存圚だず考えおいたす。

採甚掻動においおも、メルカリのミッション・バリュヌに共感する、様々なバックグラりンドの方にゞョむンしおいただけるよう、幎霢、性別、性的指向、人皮、宗教、身䜓胜力、その他蚘号に基づくあらゆる差別をなくすこずを玄束したす。

詳しくは、I&D statementをご芧ください。


なお、ご応募の際にはプラむバシヌポリシヌをご確認ください。