Overview AI-Native Systems, Storage & Data Platform Architecture We are seeking a highly technical Principal Software Engineer to drive the next generation of AI-native architecture across critical storage, database, and distributed system components. This role will
Overview We are looking for a highly motivated and technically strong Software Engineer who is passionate about building intelligent, scalable, and resilient systems at the heart of Microsoft 365’s infrastructure. You will join the Substrate Unified
Overview Microsoft’s Observability & Intelligent Cloud team is reimagining how large-scale cloud services stay reliable. We build intelligent systems that help prevent incidents, detect failures earlier, accelerate diagnosis and mitigation, and reduce operational effort across Microsoft
At Jabil (NYSE: JBL), we are proud to be a trusted partner for the worlds top brands, offering comprehensive engineering, supply chain, and manufacturing solutions. With 60 years of experience across industries and a vast network
At Jabil (NYSE: JBL), we are proud to be a trusted partner for the worlds top brands, offering comprehensive engineering, supply chain, and manufacturing solutions. With 60 years of experience across industries and a vast network
Scale-up交换机架构师 北京、上海、杭州、中国香港 全职 互联网 / 电子 / 网游 - 研发 职位描述 1、关注UALink、NVLink及高速Scale-up互联技术,理解AI训练、推理和集合通信对低时延、高带宽、确定性、扩展性和高可靠的系统需求;2、定义交换芯片端到端数据通路,包括端口与链路层、Ingress/Egress、路由、仲裁、缓存、Crossbar、VC/Credit、Ordering、Replay、QoS及多播/集合通信机制;3、建立面向真实AI通信模式的性能和流量模型,量化吞吐、时延、缓存、拥塞、扩展规模及PPA约束,指导关键微架构取舍;4、驱动架构规格和验证方案落地,协同ASIC、验证、固件、软件、软件、SerDes/PHY及系统团队完成前后仿真、芯片Bring-up和整机性能调优。 职位要求 1、交换与Fabric微架构能力:深入理解Packet/Flit/Transaction处理,能够定义端口、Ingress/Egress、路由、仲裁、共享/分布式缓存、Crossbar、调度、Traffic Management、QoS及多播数据通路;2、协议、流控与可靠性能力:掌握Credit/Backpressure、Virtual Channel、Ordering、Retry/Replay、CRC/FEC.错误隔离、死锁/活锁规避及拥塞控制,理解链路训练和SerDes/PHY接口边界;3、系统性能与芯片交付能力:能够基于AI流量进行延迟、带宽、缓存、拥塞和扩展性建模,完成时序、面积、功耗权衡;具备架构验证、Telemetry/RAS、安全隔离以及前后硅性能能调试能力。关键经历1、学历背景:微电子、电子工程、计算机、通信等相关专业,具备计算机体系结构、数字电路和网络互联基础;2、岗位经历:具备8年以上交换ASIC、NoC、GPU/CPU互联或高性能Fabric研发经历,承担过整芯片或路由、仲裁、缓存、Crossbar、流控等关键模块架构工作;3、成功经验:主导或核心参与过至少一款大型交换或互联芯片从架构、性能模型、设计验证到流片、Bring-up的完整闭环:且备UAI ink NVI infiniBand或Fthernet交换经验优先 投递...
Scale-up交换机专家 北京、上海、杭州、中国香港 全职 互联网 / 电子 / 网游 - 研发 职位描述 1、在总体架构和芯片规格指导下,负责Scale-up交换芯片关键子系统或模块的微架构设计与研发交付,包括Ingress/Egress、Packet/Flit处理、路由、仲裁、缓存、Crossbar、VC/Credit及端口数据通路等;2、负责Ordering、Retry/Replay、CRC/FEC、QoS、多播或集合通信信等协议与可靠性功能的设计实现,完成异常场景、流控依赖和死锁/活锁问题分析;3、开展模块级性能与流量分析,评估吞吐、时延、缓存占用、拥塞及PPA约束,针对All-to-All、Incast和集合通信等典型AI流量进行优化;4、协同验证、固件、软件、SerDes/PHY及系统团队,完成模块验证王、子系统集成、芯片Bring-up和整机环境下的性能与可靠性问题定位。 职位要求 1、交换数据通路研发能力:熟悉Packet/Flit处理、Ingress/Egress、路由、仲裁、缓存、Crossbar和调度机制,能够独立承担一个或多个关键模块的微架构设计、实现与验证;2、协议、流控与可靠性能力:理解Credit/Backpressure、Virtual Channel、Ordering、Retry/Replay、CRC/FECT拥塞控制机制,能够分析流控依赖、死锁和异常恢复问题;3、ASIC实现与性能调试能力:具备RTL设计、架构验证或性能建模能力,能够完成模块级时序、面积和功耗优化;熟悉Telemetry/RAS、链路调试或前后硅性能分析中的一种或多种能力。关键经历1、学历背景:微电子、电子工程、计算机、通信等相关专业,具备计算机体系结构、数字电路和网络互联基础2、岗位经历:具备5年以上交换ASIC、NoC、GPU/CPU互联或高性能Fabric研发经历,独立承担过路由、仲裁、缓存、Crossbar、流控或链路数据通路中的一个或多个关键模块决;3、成功经验:参与过至少一款交换或互联芯片从规格、设计验证到流片、Bring-up的研发过程,并完成过所负责模块的功能、性能或可靠性问题闭环;具备UALink、NVLink、InfiniBand或Ethernet相关经验优先 投递...
At Air Products, we reimagine what’s possible. By tapping into the motivation of our people and our collective experience, we create the ideas and innovations that drive us forward. When we come together – where every
We are looking for a networking test engineer with strong system‑level debugging skills to join our End‑to‑End Verification team. You will work on cutting‑edge Ethernet‑based AI clusters, owning complex issues across hardware, system software and AI
Some careers have more impact than others. If you’re looking for a career where you can make a real impression, join HSBC and discover how valued you’ll be. We are currently seeking an experienced professional to
Some careers have more impact than others. If you’re looking for a career where you can make a real impression, join HSBC and discover how valued you’ll be. We are currently seeking an experienced professional to
Location: Tokyo, Japan(On-board in China) Scopelys Security team is looking for a Sr Client Security Engineer to build a new client security function in China. You will design and ship client/server-side protections for large-scale mobile games,
Overview Windows Cloud Experiences(WCX) is developing a best-in-class virtual desktop infrastructure (VDI) solution to enable millions of customers to work with physical and virtual PCs. WCX is hiring a Senior Software Engineer for the Windows 365
Duracell is seeking an accountant to join its GBS Asia Accounting team to execute the month-end close of all financial statements, reporting and internal controls. This role will be responsible to lead, improve and execute: the
Some careers have more impact than others. If you’re looking for a career where you can make a real impression, join HSBC and discover how valued you’ll be. We are currently seeking an experienced professional to
SummaryDo you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple
Mission The mission of Speechify is to make sure that reading is never a barrier to learning. Over 50 million people use Speechify’s text-to-speech products to turn whatever they’re reading – PDFs, books, Google Docs, news
Some careers have more impact than others. If you’re looking for a career where you can make a real impression, join HSBC and discover how valued you’ll be. We are currently seeking an experienced professional to
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential
你将负责什么 1. 定义 Kimi 业务的稳定性 「稳」是什么、稳到什么程度、怎么衡量——在 Kimi,这些问题的答案由你来写。Kimi 业务正在高速发展,模型高频发布、产品形态快速演进:你不只是在维护一套静态系统,而是深入快速变化的 AI 架构与产品演进路线,在真正的前沿场景里定义可靠性。 与研发团队共建 SLO / SLI 体系,在可用性、延迟与迭代速度之间取得平衡,让可靠性目标与业务影响直接挂钩。 建立信号清晰的告警体系:分级、降噪,实现故障的分钟级发现与精准定位。 主导故障响应与复盘:快速恢复、彻底复盘、系统性改进——让团队从每一次故障中学到东西,确保同类问题不重复发生。 2. 发布工程与变更安全 为高频模型发布与产品迭代设计安全变更框架:灰度、金丝雀、自动回滚、变更可观测性,把「变更导致故障」的概率降至最低。 护航关键发布:发布前 readiness 评估(容量、依赖、回滚路径),发布中值守,发布后复盘。 把发布 runbook 沉淀为工具与流水线,把手工核验变成持续校验——每一次发布,都让 checklist 更短。 3. 可观测性与可靠性工程 与研发共建统一的 Telemetry 标准(Metrics / Logs / Traces),构建从业务指标到基础设施指标的全链路可观测性。