SOFTC Observability & AIOps

Product direction · Internal incubation

Exploring understandable, reasoning-aware, and actionable operations through unified telemetry, semantic correlation, and operations agents.

This product is in internal incubation; this page presents SOFTC’s product vision, capability boundaries, and evolution direction.

01
Why it matters

Turn fragmented signals into explainable operational context

Complex IT estates produce fragmented data, noisy alerts, and missing business semantics. This product direction uses unified data, topology, and evidence chains to support analysis and collaboration.

Core value

Move from seeing data to understanding systems and coordinating action

01

Unified observability

Align metrics, logs, traces, profiles, traffic, and events.

02

Explainable correlation

Organize incident evidence through service topology and business semantics.

03

Bounded intelligence

Rules, small models, and topology algorithms make technical judgments; LLMs explain, retrieve, interact, and assist decisions.

Capabilities

Nine capabilities across observability, reliability governance, and intelligent collaboration

01

Observability system and data model

Plan observed objects, metric definitions, labels, instrumentation, quality gates, and role-based views.

Give teams a shared operational language.
02

Unified full-stack telemetry

Collect metrics, logs, traces, profiles, traffic, events, and changes while governing transport, storage, and retention.

Create a reliable and scalable telemetry foundation.
03

Application and business observability

Connect application performance, traces, errors, versions, transactions, and business KPIs.

Move from system health to experience and business health.
04

Infrastructure and cloud-native observability

Observe hosts, networks, clouds, Kubernetes, databases, and middleware in one operating context.

Locate cross-layer bottlenecks in one health view.
05

Semantic topology and business map

Connect services, resources, calls, changes, environments, ownership, and business relationships into dynamic topology.

Turn isolated records into a reasoning-ready relationship network.
06

Event and reliability governance

Unify alerts, deduplication, suppression, aggregation, routing, SLOs, error budgets, ITSM, and reviews.

Converge large alert volumes into actionable incidents.
07

AIOps anomaly and root-cause analysis

Apply dynamic baselines, multi-signal anomalies, log clustering, topology propagation, change correlation, and root-cause ranking.

Move from manual investigation to ranked candidates and evidence.
08

Capacity, performance, and FinOps optimization

Analyze capacity trends, utilization, degradation, SLO risk, cost change, and scaling timing.

Move from reactive scaling to proactive risk and cost governance.
09

LLM enhancement and operations agents

Provide evidence-grounded Q&A, incident analysis, release risk, reporting, knowledge retrieval, and controlled automation.

Make technical judgment understandable, reviewable, and collaborative.

Architecture

A five-layer operating system from unified data to cognitive augmentation

  1. 01

    Unified operations and interaction layer

    Bring health, incidents, analysis, and collaboration to development, operations, and management roles.

    • Role workspaces
    • Service health
    • Incident collaboration
    • Operations agents
  2. 02

    AIOps and cognitive augmentation layer

    Support decisions with anomaly detection, correlation, root-cause candidates, forecasting, and LLM explanation.

    • Dynamic anomalies
    • Root-cause candidates
    • Risk forecasting
    • Knowledge and RAG
  3. 03

    Semantic correlation and reliability layer

    Organize objects, topology, incidents, changes, SLOs, business impact, and ownership as an evidence network.

    • Object model
    • Dynamic topology
    • Incidents and SLOs
    • Change and business context
  4. 04

    Full-stack observability data layer

    Collect, clean, align, standardize, transport, and store multimodal operational data.

    • Metrics and logs
    • Traces and profiles
    • Traffic and events
    • Changes and configuration
  5. 05

    Runtime objects and open integration layer

    Connect applications, cloud-native platforms, infrastructure, databases, middleware, and enterprise toolchains.

    • Applications and business
    • Kubernetes and cloud
    • Infrastructure and middleware
    • DevOps / CMDB / ITSM

Reliability loop

A continuous path from anomaly detection to organizational learning

  1. 01Detect

    Detect

    Identify anomalies from dynamic baselines, SLOs, and multimodal telemetry.

  2. 02Converge

    Converge

    Deduplicate, suppress, and aggregate alerts into actionable incidents.

  3. 03Correlate

    Correlate

    Connect topology, change, environment, ownership, and business context.

  4. 04Diagnose

    Diagnose

    Rank root-cause candidates and evidence with rules, models, and topology algorithms.

  5. 05Respond

    Respond

    Recommend runbooks and options, then coordinate execution under access, approval, and audit controls.

  6. 06Learn

    Learn

    Feed incidents, reviews, knowledge, and feedback into continuous model and rule improvement.

Key scenarios

Four operating scenarios across applications, platforms, incidents, and capacity

01

Application performance and business health

Connect application signals, traces, releases, and business KPIs to explain impact on experience and critical services.

02

Cloud-native and middleware operations

Observe clusters, containers, networks, databases, and middleware while tracing cross-layer bottlenecks through topology.

03

Incident convergence and root-cause candidates

Converge alerts, correlate topology and changes, and produce evidence-backed candidate causes and response options.

04

Capacity, SLO risk, and FinOps

Use forecasting, efficiency, and error budgets to identify capacity, reliability, and cost risks.

LLM augmentation

Make existing judgments easier to understand, reuse, and act on

LLMs do not bypass the AIOps technical judgment chain or guess anomalies and root causes; they explain, retrieve, collaborate, and support controlled execution from rule-, model-, topology-, and evidence-based results.

Operations Q&A

Answer operational questions with cited metrics, logs, topology, and incident history.

Incident analysis agent

Explain root-cause candidates, blast radius, and evidence without guessing the technical judgment.

Release risk agent

Compare changes with historical incidents, operating baselines, and dependency topology.

Operations report agent

Turn operating data into traceable briefings, reports, and incident reviews.

Knowledge and runbooks

Retrieve versioned procedures and experience inside the current incident context.

Controlled automation

Convert recommendations into approved, reversible, and auditable execution flows.

Evolution path

Build progressively around applications and measurable operating scenarios

  1. 01Unify data

    Establish observability data and event models.

  2. 02Build correlation

    Add service topology, business semantics, and change context.

  3. 03Enhance analysis

    Introduce rules, small models, and topology algorithms for technical judgment.

  4. 04Govern reliability

    Unify incidents, SLOs, changes, and collaborative response.

  5. 05Augment cognition

    Use LLMs and agents for explanation, retrieval, interaction, decisions, and controlled execution.