Hello,
I have an SNM v 4.5 system (installed on Windows Server 2003) to monitor several Windows servers. It seems to be working well and data is being collected in the graphs, etc. but the email alert messages all show a single value in the actual: field for each type of alert.
Examples:
SNM status = DOWN for:
Device (IP) : pblutl01 (10.138.128.28)
Interface : 2.67.58(Volume C avg disk queue)
Content : actual:7.596825 gt:200:5 threshold:200
Count : 5
Threshold : 5
Date & Time : Mon Mar 8 15:04:11 2010
SNM status = DOWN for:
Device (IP) : pblutl01 (10.138.128.28)
Interface : 2.67.58(Volume C avg disk queue)
Content : actual:7.596825 gt:200:5 threshold:200
Count : 5
Threshold : 5
Date & Time : Sat Mar 6 00:04:10 2010
Each monitored condition shows a different number, but the same number each time that condition is alerted. Correct values appear for each alert on the Alerts web page; only the email is incorrect.
I have tried deleting the .RRD files for one system but this had no effect. Where else might these values be stored? Is there a template file for the email alerts which is possibly corrupt?
I am out of ideas, any suggestions would be appreciated!
JDC
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
Anonymous
-
2010-03-16
More information:
I searched my folder of email alerts; it appears the actual: number does change over time for a given host and monitored attribute, but I can't see any particular pattern:
It will be '-' for a week or two, then 688.455566677 for a few days, then 0 for a week, etc.
I will do a reboot of the server SNM is installed on , and see if the 'stuck' numbers change at all.
Hope this helps…
Thanks,
JDC
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
I hadn't been receiving many alerts aside from lack of response caused by network or power outages in a while so I hadn't noticed this. However, I am getting the issue as well:
SNM status = UP for:
Device (IP) : FTR_IDF2_3560_1 (10.6.1.31)
Interface : 3.1005(Temperature )
Content : actual:47 gt:58:1 threshold:58
Date & Time : Thu Apr 22 07:15:02 2010
SNM status = DOWN for:
Device (IP) : DO_router (10.7.1.1)
Interface : 3.3(Intake )
Content : actual:24.0047744444444 gt:45:1 threshold:45
Count : 1
Threshold : 1
Date & Time : Mon Apr 12 03:15:00 2010
This is installed on a Windows server running 4.50. The only installation I have with linux is running on my home server, and I haven't received any alerts recently to test.
Thanks
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
Apologies for the late response.
Just to confirm, can you advise what is happening:
For Device : FTR_IDF2_3560_1 the 'UP' suggests 'alert=1U:alertid' is set, i.e. U = send alert when status returns to UP.
Is that correct? Is that working to your expectations?
For Device : DO_router the 'DOWN' with actual:24… gt:45:1 threshold:45 suggests there is a problem as it should only alert when greater than 45. Is that correct? Is this NOT working to your expectations?
Above suggests you are getting false positives, but you state you are not recieving alerts.
Can you run snm in terst mode to see if it 'sees' the email server?
regards
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
Sorry, ignore the "UP" alert, I had copied the wrong one. For the "DOWN" alert, the problem is that the email has the wrong number for "actual:". I am getting email alerts, and they are not false positives. The snm web interface will show the correct information for the alert. Also, that SNMP query for that device will never return decimal numbers, only whole numbers. My guess would be that the email alerting code is using an incorrect variable somewhere.
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
Anonymous
-
2011-01-04
Any suggestions on solving this one? Anyone? Anyone?
JDC
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
I'll try to describe the issue in as much detail as possible so hopefully the developer or someone with Perl skills could patch it. In both the web console under Alerts and in email Alerts, the value for "actual" is wildly incorrect. It appears that it is displaying the wrong variable. I'm not sure what this variable is. An example alert I received:
SNM status = DOWN for:
Device (IP) : DO_router (10.7.1.1)
Interface : 3.3(Intake )
Content : actual:14.9982322222222 gt:35:1 threshold:35
Count : 1
Threshold : 1
Date & Time : Mon Jan 10 13:30:00 2011
The actual temperature is currently 35 according to the graph page for this device, and had a maximum of 39, so the alert was warranted. However, the alert email and alert page on the web interface is showing "14.9982322222222" for the actual. The snmp OID being monitored for this always reports a whole number, so it is impossible that the temperature is "14.9982322222222", not to mention that the graph page is reporting the correct temperature.
Now I'm off to try and get the A/C situation fixed…
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
I'd also like to add, to JDC, the issue is consistent across other installations, I do not believe this is a corruption or other issue with our configs. I believe this is a bug with the code for either SNM or one of the Perl dependencies it is using.
If you would like to refer to this comment somewhere else in this project, copy and paste the following link:
Hello,
I have an SNM v 4.5 system (installed on Windows Server 2003) to monitor several Windows servers. It seems to be working well and data is being collected in the graphs, etc. but the email alert messages all show a single value in the actual: field for each type of alert.
Examples:
SNM status = DOWN for:
Device (IP) : pblutl01 (10.138.128.28)
Interface : 2.67.58(Volume C avg disk queue)
Content : actual:7.596825 gt:200:5 threshold:200
Count : 5
Threshold : 5
Date & Time : Mon Mar 8 15:04:11 2010
SNM status = DOWN for:
Device (IP) : pblutl01 (10.138.128.28)
Interface : 2.67.58(Volume C avg disk queue)
Content : actual:7.596825 gt:200:5 threshold:200
Count : 5
Threshold : 5
Date & Time : Sat Mar 6 00:04:10 2010
Each monitored condition shows a different number, but the same number each time that condition is alerted. Correct values appear for each alert on the Alerts web page; only the email is incorrect.
I have tried deleting the .RRD files for one system but this had no effect. Where else might these values be stored? Is there a template file for the email alerts which is possibly corrupt?
I am out of ideas, any suggestions would be appreciated!
JDC
I will try to replicate and resolve.
regards
Thomas
More information:
I searched my folder of email alerts; it appears the actual: number does change over time for a given host and monitored attribute, but I can't see any particular pattern:
It will be '-' for a week or two, then 688.455566677 for a few days, then 0 for a week, etc.
I will do a reboot of the server SNM is installed on , and see if the 'stuck' numbers change at all.
Hope this helps…
Thanks,
JDC
I hadn't been receiving many alerts aside from lack of response caused by network or power outages in a while so I hadn't noticed this. However, I am getting the issue as well:
SNM status = UP for:
Device (IP) : FTR_IDF2_3560_1 (10.6.1.31)
Interface : 3.1005(Temperature )
Content : actual:47 gt:58:1 threshold:58
Date & Time : Thu Apr 22 07:15:02 2010
SNM status = DOWN for:
Device (IP) : DO_router (10.7.1.1)
Interface : 3.3(Intake )
Content : actual:24.0047744444444 gt:45:1 threshold:45
Count : 1
Threshold : 1
Date & Time : Mon Apr 12 03:15:00 2010
This is installed on a Windows server running 4.50. The only installation I have with linux is running on my home server, and I haven't received any alerts recently to test.
Thanks
Apologies for the late response.
Just to confirm, can you advise what is happening:
For Device : FTR_IDF2_3560_1 the 'UP' suggests 'alert=1U:alertid' is set, i.e. U = send alert when status returns to UP.
Is that correct? Is that working to your expectations?
For Device : DO_router the 'DOWN' with actual:24… gt:45:1 threshold:45 suggests there is a problem as it should only alert when greater than 45. Is that correct? Is this NOT working to your expectations?
Above suggests you are getting false positives, but you state you are not recieving alerts.
Can you run snm in terst mode to see if it 'sees' the email server?
regards
Sorry, ignore the "UP" alert, I had copied the wrong one. For the "DOWN" alert, the problem is that the email has the wrong number for "actual:". I am getting email alerts, and they are not false positives. The snm web interface will show the correct information for the alert. Also, that SNMP query for that device will never return decimal numbers, only whole numbers. My guess would be that the email alerting code is using an incorrect variable somewhere.
Any suggestions on solving this one? Anyone? Anyone?
JDC
I'll try to describe the issue in as much detail as possible so hopefully the developer or someone with Perl skills could patch it. In both the web console under Alerts and in email Alerts, the value for "actual" is wildly incorrect. It appears that it is displaying the wrong variable. I'm not sure what this variable is. An example alert I received:
The actual temperature is currently 35 according to the graph page for this device, and had a maximum of 39, so the alert was warranted. However, the alert email and alert page on the web interface is showing "14.9982322222222" for the actual. The snmp OID being monitored for this always reports a whole number, so it is impossible that the temperature is "14.9982322222222", not to mention that the graph page is reporting the correct temperature.
Now I'm off to try and get the A/C situation fixed…
I'd also like to add, to JDC, the issue is consistent across other installations, I do not believe this is a corruption or other issue with our configs. I believe this is a bug with the code for either SNM or one of the Perl dependencies it is using.
I submitted a bug to the bugtracker:
https://sourceforge.net/tracker/?func=detail&aid=3154621&group_id=110780&atid=657424