Customer story · Home and building products manufacturing

JOMOO: Observability and Intelligent Operations

九牧厨卫将可观测从分散监控工具提升为企业级运行数据与决策体系,用统一语义连接资源、应用、服务、业务交易、事件和变更。
ObservabilityAIOpsIntelligent operations
01

观测数据

指标·日志·链路
02

运行中枢

事件中心
03

分析方式

语义关联
04

运维演进

AIOps

Background

Expanding applications and infrastructure produced operational data across separate tools and teams.

Core challenge

Incident detection, impact analysis, and root-cause investigation lacked shared data and semantic context.

Customer context and needs

Clarify the problems before defining the program

Pain points

应用运行状态不可见

监控主要集中在操作系统、虚拟机和部分数据库,应用、接口、方法和业务动作缺少统一观测。

告警分散且噪声高

AlertManager、夜莺等多个系统分别运行,重复告警、孤岛告警和责任不清导致值班负担较高。

多源数据缺少语义关联

指标、日志和调用链虽然存在,但标签、时间、对象和 TraceID 缺少统一标准,数据难以自动关联。

故障分析高度依赖人工

运维和开发需要在多个工具间手工拼接证据,无法快速判断影响范围、根因候选和变更关联。

Requirements

建立全域可观测数据底座

覆盖数据中心、公有云、虚拟机、容器、数据库、中间件、应用和调用链,消除数据盲区。

统一事件与稳定性治理

将多源告警去重、抑制、聚类为可处理事件,通过 SLO/SLA、分级响应和 ITSM 联动形成闭环。

建设语义数据网络

建立统一观测对象模型,关联指标、日志、链路、变更、事件和业务影响。

落地 AIOps 与认知增强

建设动态异常检测、告警聚类、变更归因、根因候选和容量分析,再由大模型负责解释、问答和决策辅助。

Goals

Translate business needs into verifiable goals

  1. 01Unify telemetry

    覆盖数据中心、公有云、虚拟机、容器、数据库、中间件、应用和调用链,消除数据盲区。

  2. 02Connect business and technical context

    将多源告警去重、抑制、聚类为可处理事件,通过 SLO/SLA、分级响应和 ITSM 联动形成闭环。

  3. 03Improve incident analysis

    建立统一观测对象模型,关联指标、日志、链路、变更、事件和业务影响。

Solution

An observability and AIOps system spanning metrics, logs, traces, topology, and events.

01

统一数据采集与存储

通过 Prometheus、OpenTelemetry、日志采集、Kafka/Flink 处理与分层存储,建立稳定的数据供应链。

02

事件告警与 SLO

建设统一告警模型、事件生命周期、P1/P2/P3 分级响应、变更关联、错误预算与复盘机制。

03

语义关联与 AIOps

建设服务依赖拓扑、上下文时间窗口、影响模型、异常检测、时间聚合、根因候选和自学习模型。

04

大模型增强与流程集成

通过运维知识库、故障分析智能体、发布风险智能体和报告智能体,连接 DevOps、ITSM、CMDB 和 IAM。

Key architecture

统一可观测与 AIOps 能力架构

从全栈数据采集到语义关联、智能分析和大模型认知增强,形成运行闭环。
01

应用领域与观测对象

将运行数据对齐到业务

制造领域研发领域供应链营销领域业务交易服务 / API
02

统一可观测数据底座

统一采集、标准、传输和存储

指标日志调用链变更数据统一标签分层存储
03

事件与语义关联

把原始数据转化为可理解上下文

告警降噪事件模型观测对象服务拓扑业务影响SLO / SLA
04

AIOps 智能分析

围绕实际运维场景输出判断

动态异常检测告警聚类变更归因根因候选影响评估容量与 FinOps
05

大模型增强与流程闭环

解释、检索、交互与受治理执行

智能运维问答故障分析智能体发布风险智能体运维报告ITSMDevOps / CMDB / IAM

Implementation journey

A phased path from assessment to sustainable operations

  1. 01
    Plan data and scenarios

    Define scenarios, data models, service identities, business semantics, and the roadmap.

  2. 02
    Unify observability and semantic correlation

    Build collection standards, shared data, service topology, business maps, and event systems.

  3. 03
    Add AIOps and intelligent collaboration

    Introduce anomaly detection, correlation, and governed operations agents on shared evidence.

Implementation principle

The program begins with data models and collection standards, then adds correlation and intelligent operations capabilities.

Results

Outcomes across efficiency, quality, resilience, and operations

01

故障定位路径更短

开发和运维可以围绕服务、接口和业务链路查看指标、日志、调用链、变更和历史事件。

02

告警转换为可治理事件

同一故障引发的重复告警通过去重、抑制、聚类和拓扑关联收敛为事件,降低值班认知负担。

03

稳定性从经验转向指标闭环

通过服务健康、SLO、关键链路、事件闭环和复盘改进,让稳定性可度量、可跟踪、可持续改进。

04

运维知识可复用

历史故障、根因证据、处置步骤和复盘结论进入知识库,为智能问答和决策辅助提供可追溯依据。

Future evolution

Capabilities continue to evolve after platform launch

From observability to evidence-driven intelligent operations

Rules, small models, and topology algorithms make technical judgments; LLMs explain, retrieve, interact, and assist decisions.

Capability tags

ObservabilityAIOpsIntelligent operations