观测数据
指标·日志·链路Customer story · Home and building products manufacturing
JOMOO: Observability and Intelligent Operations
九牧厨卫将可观测从分散监控工具提升为企业级运行数据与决策体系,用统一语义连接资源、应用、服务、业务交易、事件和变更。运行中枢
事件中心分析方式
语义关联运维演进
AIOpsBackground
Expanding applications and infrastructure produced operational data across separate tools and teams.
Core challenge
Incident detection, impact analysis, and root-cause investigation lacked shared data and semantic context.
Customer context and needs
Clarify the problems before defining the program
Pain points
应用运行状态不可见
监控主要集中在操作系统、虚拟机和部分数据库,应用、接口、方法和业务动作缺少统一观测。
告警分散且噪声高
AlertManager、夜莺等多个系统分别运行,重复告警、孤岛告警和责任不清导致值班负担较高。
多源数据缺少语义关联
指标、日志和调用链虽然存在,但标签、时间、对象和 TraceID 缺少统一标准,数据难以自动关联。
故障分析高度依赖人工
运维和开发需要在多个工具间手工拼接证据,无法快速判断影响范围、根因候选和变更关联。
Requirements
建立全域可观测数据底座
覆盖数据中心、公有云、虚拟机、容器、数据库、中间件、应用和调用链,消除数据盲区。
统一事件与稳定性治理
将多源告警去重、抑制、聚类为可处理事件,通过 SLO/SLA、分级响应和 ITSM 联动形成闭环。
建设语义数据网络
建立统一观测对象模型,关联指标、日志、链路、变更、事件和业务影响。
落地 AIOps 与认知增强
建设动态异常检测、告警聚类、变更归因、根因候选和容量分析,再由大模型负责解释、问答和决策辅助。
Goals
Translate business needs into verifiable goals
- 01Unify telemetry
覆盖数据中心、公有云、虚拟机、容器、数据库、中间件、应用和调用链,消除数据盲区。
- 02Connect business and technical context
将多源告警去重、抑制、聚类为可处理事件,通过 SLO/SLA、分级响应和 ITSM 联动形成闭环。
- 03Improve incident analysis
建立统一观测对象模型,关联指标、日志、链路、变更、事件和业务影响。
Solution
An observability and AIOps system spanning metrics, logs, traces, topology, and events.
统一数据采集与存储
通过 Prometheus、OpenTelemetry、日志采集、Kafka/Flink 处理与分层存储,建立稳定的数据供应链。
事件告警与 SLO
建设统一告警模型、事件生命周期、P1/P2/P3 分级响应、变更关联、错误预算与复盘机制。
语义关联与 AIOps
建设服务依赖拓扑、上下文时间窗口、影响模型、异常检测、时间聚合、根因候选和自学习模型。
大模型增强与流程集成
通过运维知识库、故障分析智能体、发布风险智能体和报告智能体,连接 DevOps、ITSM、CMDB 和 IAM。
Key architecture
统一可观测与 AIOps 能力架构
从全栈数据采集到语义关联、智能分析和大模型认知增强,形成运行闭环。应用领域与观测对象
将运行数据对齐到业务
统一可观测数据底座
统一采集、标准、传输和存储
事件与语义关联
把原始数据转化为可理解上下文
AIOps 智能分析
围绕实际运维场景输出判断
大模型增强与流程闭环
解释、检索、交互与受治理执行
Implementation journey
A phased path from assessment to sustainable operations
- 01Plan data and scenarios
Define scenarios, data models, service identities, business semantics, and the roadmap.
- 02Unify observability and semantic correlation
Build collection standards, shared data, service topology, business maps, and event systems.
- 03Add AIOps and intelligent collaboration
Introduce anomaly detection, correlation, and governed operations agents on shared evidence.
The program begins with data models and collection standards, then adds correlation and intelligent operations capabilities.
Results
Outcomes across efficiency, quality, resilience, and operations
故障定位路径更短
开发和运维可以围绕服务、接口和业务链路查看指标、日志、调用链、变更和历史事件。
告警转换为可治理事件
同一故障引发的重复告警通过去重、抑制、聚类和拓扑关联收敛为事件,降低值班认知负担。
稳定性从经验转向指标闭环
通过服务健康、SLO、关键链路、事件闭环和复盘改进,让稳定性可度量、可跟踪、可持续改进。
运维知识可复用
历史故障、根因证据、处置步骤和复盘结论进入知识库,为智能问答和决策辅助提供可追溯依据。
SOFTC products and services
Future evolution
Capabilities continue to evolve after platform launch
From observability to evidence-driven intelligent operations
Rules, small models, and topology algorithms make technical judgments; LLMs explain, retrieve, interact, and assist decisions.
Capability tags