Unified observability
Align metrics, logs, traces, profiles, traffic, and events.
Product direction · Internal incubation
Exploring understandable, reasoning-aware, and actionable operations through unified telemetry, semantic correlation, and operations agents.
This product is in internal incubation; this page presents SOFTC’s product vision, capability boundaries, and evolution direction.
01
Why it matters
Complex IT estates produce fragmented data, noisy alerts, and missing business semantics. This product direction uses unified data, topology, and evidence chains to support analysis and collaboration.
Core value
Align metrics, logs, traces, profiles, traffic, and events.
Organize incident evidence through service topology and business semantics.
Rules, small models, and topology algorithms make technical judgments; LLMs explain, retrieve, interact, and assist decisions.
Capabilities
Plan observed objects, metric definitions, labels, instrumentation, quality gates, and role-based views.
Give teams a shared operational language.Collect metrics, logs, traces, profiles, traffic, events, and changes while governing transport, storage, and retention.
Create a reliable and scalable telemetry foundation.Connect application performance, traces, errors, versions, transactions, and business KPIs.
Move from system health to experience and business health.Observe hosts, networks, clouds, Kubernetes, databases, and middleware in one operating context.
Locate cross-layer bottlenecks in one health view.Connect services, resources, calls, changes, environments, ownership, and business relationships into dynamic topology.
Turn isolated records into a reasoning-ready relationship network.Unify alerts, deduplication, suppression, aggregation, routing, SLOs, error budgets, ITSM, and reviews.
Converge large alert volumes into actionable incidents.Apply dynamic baselines, multi-signal anomalies, log clustering, topology propagation, change correlation, and root-cause ranking.
Move from manual investigation to ranked candidates and evidence.Analyze capacity trends, utilization, degradation, SLO risk, cost change, and scaling timing.
Move from reactive scaling to proactive risk and cost governance.Provide evidence-grounded Q&A, incident analysis, release risk, reporting, knowledge retrieval, and controlled automation.
Make technical judgment understandable, reviewable, and collaborative.Architecture
Bring health, incidents, analysis, and collaboration to development, operations, and management roles.
Support decisions with anomaly detection, correlation, root-cause candidates, forecasting, and LLM explanation.
Organize objects, topology, incidents, changes, SLOs, business impact, and ownership as an evidence network.
Collect, clean, align, standardize, transport, and store multimodal operational data.
Connect applications, cloud-native platforms, infrastructure, databases, middleware, and enterprise toolchains.
Reliability loop
Identify anomalies from dynamic baselines, SLOs, and multimodal telemetry.
Deduplicate, suppress, and aggregate alerts into actionable incidents.
Connect topology, change, environment, ownership, and business context.
Rank root-cause candidates and evidence with rules, models, and topology algorithms.
Recommend runbooks and options, then coordinate execution under access, approval, and audit controls.
Feed incidents, reviews, knowledge, and feedback into continuous model and rule improvement.
Key scenarios
Connect application signals, traces, releases, and business KPIs to explain impact on experience and critical services.
Observe clusters, containers, networks, databases, and middleware while tracing cross-layer bottlenecks through topology.
Converge alerts, correlate topology and changes, and produce evidence-backed candidate causes and response options.
Use forecasting, efficiency, and error budgets to identify capacity, reliability, and cost risks.
LLM augmentation
LLMs do not bypass the AIOps technical judgment chain or guess anomalies and root causes; they explain, retrieve, collaborate, and support controlled execution from rule-, model-, topology-, and evidence-based results.
Answer operational questions with cited metrics, logs, topology, and incident history.
Explain root-cause candidates, blast radius, and evidence without guessing the technical judgment.
Compare changes with historical incidents, operating baselines, and dependency topology.
Turn operating data into traceable briefings, reports, and incident reviews.
Retrieve versioned procedures and experience inside the current incident context.
Convert recommendations into approved, reversible, and auditable execution flows.
Evolution path
Establish observability data and event models.
Add service topology, business semantics, and change context.
Introduce rules, small models, and topology algorithms for technical judgment.
Unify incidents, SLOs, changes, and collaborative response.
Use LLMs and agents for explanation, retrieval, interaction, decisions, and controlled execution.