Menu ▾ ▴

Incorrect actual: in email alerts

Help
Anonymous
2010-03-09
2013-04-11
  • Anonymous

    Anonymous - 2010-03-09

    Hello,
         I have an SNM v 4.5 system (installed on Windows Server 2003) to monitor several Windows servers. It seems to be working well and data is being collected in the graphs, etc. but the email alert messages all show a single value in the actual: field for each type of alert.
    Examples:
    SNM status  = DOWN for:
    Device (IP) : pblutl01 (10.138.128.28)
    Interface   : 2.67.58(Volume C avg disk queue)
    Content     : actual:7.596825 gt:200:5 threshold:200
    Count       : 5
    Threshold   : 5
    Date & Time : Mon Mar  8 15:04:11 2010

    SNM status  = DOWN for:
    Device (IP) : pblutl01 (10.138.128.28)
    Interface   : 2.67.58(Volume C avg disk queue)
    Content     : actual:7.596825 gt:200:5 threshold:200
    Count       : 5
    Threshold   : 5
    Date & Time : Sat Mar  6 00:04:10 2010

    Each monitored condition shows a different number, but the same number each time that condition is alerted. Correct values appear for each alert on the Alerts web page; only the email is incorrect.

    I have tried deleting the .RRD files for one system but this had no effect. Where else might these values be stored? Is there a template file for the email alerts which is possibly corrupt?

    I am out of ideas, any suggestions would be appreciated!

    JDC

     
  • Thomas Price

    Thomas Price - 2010-03-11

    I will try to replicate and resolve.

    regards
    Thomas

     
  • Anonymous

    Anonymous - 2010-03-16

    More information:
    I searched my folder of email alerts; it appears the actual: number does change over time for a given host and monitored attribute, but I can't see any particular pattern:
    It will be '-' for a week or two, then 688.455566677 for a few days, then 0 for a week, etc.

    I will do a reboot of the server SNM is installed on , and see if the 'stuck' numbers change at all.

    Hope this helps…

    Thanks,
    JDC

     
  • Richard Moeller

    Richard Moeller - 2010-04-27

    I hadn't been receiving many alerts aside from lack of response caused by network or power outages in a while so I hadn't noticed this. However, I am getting the issue as well:

    SNM status  = UP for:
    Device (IP) : FTR_IDF2_3560_1 (10.6.1.31)
    Interface   : 3.1005(Temperature          )
    Content     : actual:47 gt:58:1 threshold:58
    Date & Time : Thu Apr 22 07:15:02 2010

    SNM status  = DOWN for:
    Device (IP) : DO_router (10.7.1.1)
    Interface   : 3.3(Intake               )
    Content     : actual:24.0047744444444 gt:45:1 threshold:45
    Count       : 1
    Threshold   : 1
    Date & Time : Mon Apr 12 03:15:00 2010

    This is installed on a Windows server running 4.50. The only installation I have with linux is running on my home server, and I haven't received any alerts recently to test.

    Thanks

     
  • Thomas Price

    Thomas Price - 2010-05-08

    Apologies for the late response.
    Just to confirm, can you advise what is happening:
    For Device : FTR_IDF2_3560_1 the 'UP' suggests 'alert=1U:alertid' is set, i.e. U = send alert when status returns to UP.
    Is that correct? Is that working to your expectations?
    For Device : DO_router the 'DOWN' with actual:24… gt:45:1 threshold:45 suggests there is a problem as it should only alert when greater than 45. Is that correct? Is this NOT working to your expectations?
    Above suggests you are getting false positives, but you state you are not recieving alerts.
    Can you run snm in terst mode to see if it 'sees' the email server?
    regards

     
  • Richard Moeller

    Richard Moeller - 2010-05-10

    Sorry, ignore the "UP" alert, I had copied the wrong one. For the "DOWN" alert, the problem is that the email has the wrong number for "actual:". I am getting email alerts, and they are not false positives. The snm web interface will show the correct information for the alert. Also, that SNMP query for that device will never return decimal numbers, only whole numbers. My guess would be that the email alerting code is using an incorrect variable somewhere.

     
  • Anonymous

    Anonymous - 2011-01-04

    Any suggestions on solving this one? Anyone? Anyone?

    JDC

     
  • Richard Moeller

    Richard Moeller - 2011-01-10

    I'll try to describe the issue in as much detail as possible so hopefully the developer or someone with Perl skills could patch it. In both the web console under Alerts and in email Alerts, the value for "actual" is wildly incorrect. It appears that it is displaying the wrong variable. I'm not sure what this variable is. An example alert I received:

    SNM status  = DOWN for:
    Device (IP) : DO_router (10.7.1.1)
    Interface   : 3.3(Intake               )
    Content     : actual:14.9982322222222 gt:35:1 threshold:35
    Count       : 1
    Threshold   : 1
    Date & Time : Mon Jan 10 13:30:00 2011

    The actual temperature is currently 35 according to the graph page for this device, and had a maximum of 39, so the alert was warranted. However, the alert email and alert page on the web interface is showing "14.9982322222222" for the actual. The snmp OID being monitored for this always reports a whole number, so it is impossible that the temperature is "14.9982322222222", not to mention that the graph page is reporting the correct temperature.

    Now I'm off to try and get the A/C situation fixed…

     
  • Richard Moeller

    Richard Moeller - 2011-01-10

    I'd also like to add, to JDC, the issue is consistent across other installations, I do not believe this is a corruption or other issue with our configs. I believe this is a bug with the code for either SNM or one of the Perl dependencies it is using.

     

Log in to post a comment.