Alerting
Questions about alerting and problem detection in Dynatrace.
cancel
Showing results for 
Show  only  | Search instead for 
Did you mean: 

Per-disk low disk space alerting with individual thresholds at scale

JoãoSilva
Visitor

We are currently trying to resolve an issue with low disk space problems in our Dynatrace environment.

Dynatrace is opening low disk space problems for a large number of discovered disks, but the client never requested alerting for all of them. We only want problems to be generated for disks that the client has explicitly asked us to monitor.

The environment contains thousands of disks, so creating individual anomaly detection overrides for every disk that should not alert does not seem scalable or maintainable.

There is also another requirement: the client can define different thresholds for different disks.

For example:

  • Disk A: alert below 5% free

  • Disk B: alert below 10% free

  • Disk C: alert below 15% free

  • Disk 😧 no low disk space alerting

Ideally, we would like an opt-in model where only explicitly configured disks are evaluated, and each disk can have its own threshold.

We would also prefer not to use tags as the configuration mechanism. While tags could probably be used to build a custom solution today, we are trying to avoid designing this around tags because of the direction Dynatrace is moving with Grail, Smartscape, newer entity models, DQL-based configuration, and newer anomaly detection capabilities. We would rather use a mechanism that is more aligned with the current and future Dynatrace platform.

We have been looking at options such as:

  • Disk Edge anomaly detection

  • Grail lookup tables

  • DQL-based anomaly detection

  • Record-based anomaly detection

  • Any newer native configuration mechanism

Conceptually, the ideal configuration would be something like:

Disk Threshold

Disk A5%
Disk B10%
Disk C15%

Any disk not present in that configuration would simply not generate a low disk space problem.

The main goal is to avoid maintaining thousands of individual Settings API objects or anomaly detection overrides while still keeping the resulting alerts correctly associated with the DISK entity.

Has anyone implemented something similar with the newer Dynatrace capabilities?

Is there currently a recommended/native approach for managing per-disk alerting scope and individual thresholds at scale?

5 REPLIES 5

AntonPineiro
DynaMight Guru
DynaMight Guru

Hi,

If you are in SaaS you can use Anomaly detector base on timeseries data. Or using disk Edge. Not sure if you can filter in/out entities using Disk Edge without metadata (base on hostname for example).

If not, the most flexible way would be anomaly detector app.

Best regards

❤️ Emacs ❤️ Vim ❤️ Bash ❤️ Perl

For thousands of disks you end up reaching the limit of dimensions available for the Anomaly detector which then impacts the alerting if i'm not mistaken

And for disks edge, doesn't really work for our case because you might want to monitor 2 disks from a particular host but not the remaining 10 from that host.

Best regards

Hi,

I am not aware of any dimension limit under Grail. I had just seen that behaviour using classic metric events, not new Anomaly Detector app.

I cannot understand which issue you had about disk edge. You can alert only about 2 disks and ignore others, if you want.

Best regards

❤️ Emacs ❤️ Vim ❤️ Bash ❤️ Perl

rastislav_danis
DynaMight Pro
DynaMight Pro

I would recommend to create environment disk edge rule for each disk and disable (if enabled) all other default disk edge rules and also disable (if enabled) all "Anomaly detection for infrastructure: Disk" switches.

Alanata a.s.

t_pawlak
DynaMight Leader
DynaMight Leader

Hi,
In my opinion, Disk Edge is probably the closest native solution at the moment.

You can disable the default disk alerting and keep rules only for disks that should actually be monitored. The challenge is the individual threshold per disk in that case, you may still end up maintaining a large number of rules.

If disks can be grouped by thresholds, for example 5%, 10%, 15%, then Disk Edge should work quite well. If every disk needs a completely independent threshold, I’m not aware of a simple native lookup-based mechanism for that today.

Featured Posts