Discover community-driven integrations, libraries, and resources to extend Datadog across your stack. Filter by platform, data type, use case, and more.
Immediately improve your systems' reliability with chaos engineering, enabling controlled simulation of turbulent conditions. The integration provides insights into system resilience and incident management.
Superwise is an advanced model observability platform designed to monitor machine learning models in production. This integration provides real-time tracking of model performance, data drift, and operational metrics, enabling users to detect issues, ensure model reliability, and maintain compliance throughout the model lifecycle.
Monitor and centralize Datadog incidents with TaskCall, providing real-time incident response and automated resolution. The integration enables bi-directional incident synchronization and improved impact visibility.
The Datadog TorchServe integration enables comprehensive monitoring of your TorchServe instances by collecting metrics, events, and logs from the Inference API, Management API, and OpenMetrics endpoints. Track the overall health status, model performance, and custom metrics, and receive alerts on key events such as model additions or removals. This integration supports flexible configuration for hosts, Docker, and Kubernetes environments, helping you ensure your TorchServe deployments are performing optimally and issues are detected quickly.
Visualize metrics for databases and kafka clusters from Upstash, providing insights into serverless data services. The integration enables monitoring of Redis and Kafka operations.