R2R 远程站
网络自动化软件主管 Network Automation Software Lead
公司TensorWave
薪资面议
岗位性质全职
要求地区全球
经验要求5-10年
发布时间2026-08-05
投递方式开通会员后查看
投递链接开通会员后查看
岗位来自公开渠道,时效性以官网为准。反馈失效?
🤖 AI 速读
- 8年以上网络自动化,Python/Go扎实
- 端到端ZTP系统构建经验,GPU网络加分
- 熟悉SONiC/gNMI,有团队管理经验优先
完整 JD
岗位描述
Own the end-to-end ZTP pipeline: bare-metal switch boot → image + base config → registration in source of truth → full intended config → validation → production — with zero human intervention. 拥有端到端 ZTP 管道:裸机交换机启动 → 映像 + 基本配置 → 在事实来源中注册 → 完整的预期配置 → 验证 → 生产 - 零人工干预。 Build intent-based config generation off a network source of truth / IPAM, with GitOps-style deployment, pre/post-change validation, and safe rollout and rollback. 通过 GitOps 式部署、更改前/更改后验证以及安全推出和回滚,基于网络事实来源 / IPAM 构建基于意图的配置生成。 Establish network validation and pre-deployment testing (snapshot/digital-twin testing) so changes are caught before they hit production fabrics. 建立网络验证和预部署测试(快照/数字孪生测试),以便在变更影响生产结构之前捕获变更。 Build streaming telemetry and metrics/logging pipelines (gNMI / OpenConfig) for fabric health. 构建流式遥测和指标/日志记录管道(gNMI / OpenConfig)以确保结构健康。 Instrument what matters for GPU networks: RoCE health (PFC/ECN counters), optics and link errors, BGP / EVPN state, capacity and utilization. 检测对 GPU 网络重要的因素:RoCE 运行状况(PFC/ECN 计数器)、光学和链路错误、BGP/EVPN 状态、容量和利用率。 Deliver dashboards and alerting the network team actually uses — signal, not noise. 提供仪表板并提醒网络团队实际使用的信号,而不是噪音。 Gather requirements, build self-service APIs and interfaces, and relentlessly accelerate their deployment velocity. 收集需求,构建自助服务 API 和接口,并不断加快部署速度。 Partner closely so the tooling reflects how the network is actually operated and turned up. 密切合作,使工具能够反映网络的实际运营和运行方式。 Hire, mentor, and grow a small team of software engineers and SREs; own roadmap, prioritization, and delivery — while still carrying a meaningful share of the code yourself. 雇用、指导和发展一个由软件工程师和 SRE 组成的小团队;自己的路线图、优先级和交付——同时仍然自己承担有意义的代码份额。 Set technical direction and standards, and ensure clean integration points with the broader platform stack (infra provisioning, CI/CD, secrets, identity, existing observability). 设定技术方向和标准,并确保与更广泛的平台堆栈(基础设施配置、CI/CD、秘密、身份、现有可观察性)的清晰集成点。 Bring software engineering rigor to network automation: code review, testing, release management, and on-call ownership. 将软件工程的严谨性引入网络自动化:代码审查、测试、发布管理和待命所有权。
岗位要求
8+ years of relevant experience 8年以上相关经验 Proven experience building network automation at scale — ideally at a hyperscaler, large cloud, or large-scale datacenter / AI-infrastructure operator. 拥有大规模构建网络自动化的丰富经验——最好是在超大规模、大型云或大型数据中心/人工智能基础设施运营商。 You've built or been a core contributor to a ZTP / device-provisioning system end-to-end , not just maintained one. 您已经构建了端到端的 ZTP/设备配置系统,或者是该系统的核心贡献者,而不仅仅是维护系统。 Strong software engineering fundamentals: Python and/or Go , with real production practices (version control, testing, CI/CD, code review). 强大的软件工程基础:Python 和/或 Go,以及实际的生产实践(版本控制、测试、CI/CD、代码审查)。 Hands-on depth with datacenter Clos fabrics and the protocols that run them: BGP, EVPN/VXLAN, and ideally RoCEv2 / RDMA for GPU networks at scale. 深入实践数据中心 Clos 结构以及运行它们的协议:BGP、EVPN/VXLAN,以及适用于大规模 GPU 网络的理想 RoCEv2/RDMA。 Fluency with modern network automation tech: gNMI/gNOI, OpenConfig/YANG, NETCONF ; source-of-truth systems (NetBox / Nautobot); NOS platforms (SONiC/FRR or vendor equivalents); tooling like Nornir / NAPALM / Ansible. 熟练掌握现代网络自动化技术:gNMI/gNOI、OpenConfig/YANG、NETCONF;真实来源系统(NetBox / Nautobot); NOS 平台(SONiC/FRR 或同等供应商); Nornir / NAPALM / Ansible 等工具。 Experience with observability / telemetry pipelines (Prometheus, Grafana, Kafka, OpenTelemetry, or similar). 拥有可观测性/遥测管道(Prometheus、Grafana、Kafka、OpenTelemetry 或类似管道)方面的经验。 Comfort running services on Kubernetes / containers . 在 Kubernetes/容器上舒适地运行服务。 Leadership : you've led a team or been the clear technical owner of a platform, and you instinctively treat internal users as customers. 领导力:您领导过一个团队或者是一个平台的明确技术所有者,并且您本能地将内部用户视为客户。 Experience with GPU / AI training or inference clusters and their backend networks. 拥有 GPU/AI 训练或推理集群及其后端网络的经验。 Familiarity with the AMD networking ecosystem (Pensando DPUs, Ultra Ethernet) or building on Ethernet-based RDMA fabrics. 熟悉 AMD 网络生态系统(Pensando DPU、Ultra 以太网)或基于以太网的 RDMA 结构进行构建。 Whitebox / disaggregated networking and SONiC at scale. 大规模白盒/分解网络和 SONiC。 Network validation / digital-twin tooling (e.g., Batfish, containerlab). 网络验证/数字孪生工具(例如 Batfish、containerlab)。 Multi-site / multi-region datacenter buildouts. 多站点/多区域数据中心扩建。
💡 你可能也感兴趣
开通会员,解锁 AI 匹配 + 结构化改简历
你已看到岗位详情与 JD。想进一步看「简历真实匹配度」「能力缺口」「逐条改写建议」,并进入结构化编辑器按 JD 改简历、导出 PDF,并生成「求职联系文案」和「面试建议」?开通会员即可。
结构化改写建议 · 左侧编辑,右侧采纳
① 你的简历(工作经历 / 项目经历可编辑,其余只读)
② AI 改写建议(每条对应一个经历,可一键复制表述)
③ 实时预览
点「预览 / 打印 PDF」生成
重新分析免费;导出 PDF 需会员或单次购买(每次导出消耗 1 次改简历权益)。
生成求职联系文案
使用场景
中文版本
English version
面试备战建议
我的求职偏好
以下用于「为我匹配」按契合度排序,不等于只看这些;要精确组合筛选请去岗位库。