小红书 社会招聘
Devops平台研发工程师-国际化
美国 ·经验不限·学历不限
美国
薪资面议
前往 小红书 官网投递
职位描述
1、负责国内 DevOps、高可用平台在海外区域的部署、集成、适配和运营支持,推动平台能力按计划落地。
2、评估海外公有云、私有云或数据中心环境在计算、网络、存储、Kubernetes、安全及合规等方面的差异,制定并执行适配方案。
3、在 DevOps 方向,负责服务发布平台、变更管控及规模化运维平台的海外接入,完善 CI/CD、基础设施自动化和发布运维流程。
4、在高可用方向,负责应急预案、故障切换、容灾验证及混沌演练等能力的落地,跟进问题整改并持续提升海外业务可靠性。
5、参与海外生产环境的日常运维和故障处置,能够快速定位发布、基础设施及平台集成问题,并推动根因分析和闭环改进。
6、沉淀部署文档、操作手册、应急预案和标准作业流程,提升海外平台交付和运维效率。
7、与国内平台研发团队以及海外基础设施、网络、安全和业务团队协作,管理实施进度、依赖和风险,及时反馈海外场景的产品改进需求。
1、Drive the rollout, integration, localization, and operational support of DevOps and high-availability platforms across international regions, ensuring platform capabilities are delivered reliably and on schedule.
2、Evaluate infrastructure differences across international public cloud, private cloud, and data center environments, including compute, networking, storage, Kubernetes, security, and compliance requirements. Design and implement practical adaptation plans.
3、Support the international adoption of internal DevOps platforms, including service release systems, change management platforms, and large-scale operations tooling. Improve CI/CD pipelines, infrastructure automation, and release operations workflows.
4、Implement high-availability capabilities for international environments, including incident response playbooks, failover mechanisms, disaster recovery validation, and chaos engineering practices. Track remediation actions and continuously improve service reliability.
5、Participate in production operations and incident response for international environments. Quickly identify and troubleshoot issues related to releases, infrastructure, and platform integration, and drive root-cause analysis and closed-loop improvements.
6、Maintain deployment guides, operational runbooks, incident response playbooks, and standard operating procedures to improve delivery quality and operational efficiency across international regions.
7、Work closely with central platform R&D teams, as well as international infrastructure, networking, security, and business teams. Manage project progress, dependencies, and risks, and provide feedback to improve platform capabilities for international scenarios.
【任职要求】
1、本科及以上学历,计算机、软件工程或相关专业优先。
2、具备 DevOps、SRE、云平台或基础设施工程相关经验,有较强的实际交付和问题解决能力。
3、熟悉 Linux、计算机网络及常见基础设施组件,能够独立进行系统部署、配置和故障排查。
4、熟悉 Kubernetes、Helm、Terraform 等云原生及基础设施即代码技术;具有 KubeVela 使用经验者优先。
5、熟悉 Go、Java、Python 或 Shell 中的一种或多种语言,能够编写自动化工具,并具备阅读和调试服务端代码的能力。
6、熟悉至少一种主流公有云或企业级私有云环境,理解多区域部署、网络连通、权限管理及安全隔离等基本概念。
7、了解监控、日志和告警体系,具有 Prometheus、Grafana、ELK/OpenSearch 或 OpenTelemetry 等相关经验。
8、中英文流利,能够在国际化团队环境中进行技术沟通与协作。
9、责任心强,执行力和推动力良好,能够将中心平台的设计方案转化为稳定、可维护的本地实施。
1、Bachelor’s degree or above in Computer Science, Software Engineering, or a related field.
2、Hands-on experience in DevOps, SRE, cloud platforms, or infrastructure engineering, with strong delivery and troubleshooting capabilities.
3、Solid understanding of Linux, computer networking, and common infrastructure components. Able to independently perform deployment, configuration, and issue diagnosis.
4、Familiar with cloud-native and Infrastructure-as-Code technologies, such as Kubernetes, Helm, and Terraform. Experience with KubeVela is a plus.
5、Proficient in at least one programming or scripting language, such as Go, Java, Python, or Shell. Able to build automation tools and read/debug backend service code.
6、Familiar with at least one major public cloud or enterprise private cloud environment. Understand concepts such as multi-region deployment, network connectivity, access control, and security isolation.
7、Familiar with monitoring, logging, and alerting systems. Experience with Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, or similar tools is preferred.
8、Fluent in both English and Chinese, with the ability to communicate effectively in a cross-region, international team environment.
9、Strong ownership, execution, and communication skills. Able to translate central platform designs into stable, maintainable implementations in local international environments.
加分项
-有内部平台海外落地、基础设施迁移或多区域部署经验。
-有大型互联网公司 DevOps 工具链或规模化运维经验。
-有容灾、混沌工程、容量管理或重大故障处置经验。
-了解数据驻留、隐私保护及所在地区基础设施合规要求。
Preferred Qualifications
-Experience rolling out internal platforms in international regions, infrastructure migration, or multi-region deployment.
-Experience with DevOps toolchains or large-scale operations platforms in major internet or cloud companies.
-Experience in disaster recovery, chaos engineering, capacity management, or major incident response.
-Understanding of data residency, privacy protection, and regional infrastructure compliance requirements.
官网发布:2026-08-13 · 最后确认在招:2026-09-03 23:16:54 · 来源平台:xiaohongshu