Menu ▾ ▴

#157 Disk problem stops wrapper from pinging JVM

open
Misc (42)
5
2007-03-17
2007-02-14
Anonymous
No

Hello,

My JVM runs as a Windows Service with the wrapper, and it logs information to its stdout many times per second. It is the job of the Wrapper to actually write this JVM output to the disk. But the disk was unreachable for more than 30 seconds (the raid timeout happens to be 40 seconds, and in my windows event log, an event was logged that the raid timeout was exceeded at least once).

The wrapper logfile looks like this:

13/02 17:43:12:355[-Default-2] [ServerHost1](#=61,1)11478-16-Roeren doseren : pli=[running], Cmd=no_command[1]
13/02 17:43:33:527[-Default-2] Executor : execute PFC Uitsteken Lijn12 [batch=1]
JVM requested a restart.
Wrapper Process has not received any CPU time for 72 seconds. Extending timeouts.
Wrapper Manager: The Wrapper code did not ping the JVM for 30 seconds. Quit and let the Wrapper resynch.
13/02 17:44:05:542[-Default-2] [ServerHost1](#=25780,1)11469-1-Kiepen : pli=[running], Cmd=no_command[1]
JVM Process has not received any CPU time for 29 seconds. Extending timeouts.
13/02 17:44:02:980[rThread-29] server ServerHost1 is closing his services
13/02 17:44:45:652[rThread-29] Stop service VisualizationServer [prior=50] in serverHost 'ServerHost1'
13/02 17:44:45:715[-Default-2] [ServerHost1](#=25780,1)11469-1-Verdelen : pli=[running], Cmd=no_command[1]
13/02 17:44:45:715[-Default-2] [ServerHost1](#=25780,1)11469-1-Voorwalsen onder : pli=[running], Cmd=no_command[1]

Both the wrapper and JVM are complaining that they are not receiving enough CPU time, but does this really mean that the CPU of the computer was heavily in use by some other process at the time?

The JVM decided to restart because it had not been pinged for more than 30 seconds.

Why did the wrapper not ping the JVM?
1) Because the wrapper did not get any CPU to do so
2) Because the wrapper was unable to write logging to the disk, and this same thread was also responsible for pinging the JVM. In my opinion, separate threads should do the logging and the pinging.

If cause 2 is the case, then separate threads should be used for logging and pinging.

Discussion

  • Leif Mortenson

    Leif Mortenson - 2007-03-17

    Logged In: YES
    user_id=228081
    Originator: NO

    The problem is close to #2. The Wrapper's main loop is responsible for pinging the JVM, logging its console output and several other tasks. A second, very simple thread, is responsible for counting cycles and works as a timer that is independent of the system clock.

    In this case, the main thread hung up while attempting to write output to your disk because it was blocking. That blocking period appears to have been 72 seconds. Once it continued, it noticed that 72 seconds worth of ticks had passed on a single cyle of the main loop and logged an appropriate message.

    The Java process has a failsafe built in to cause it to shut itself down after a period of not being pinged to avoid it becoming a zombie process. That time period can be extended by changing the wrapper.ping.timeout property.

    What caused this long freezing of your RAID disc? It seems like most applications would have the same problem. Extending the ping timeout would prevent the JVM from restarting itself.

    Cheers,
    Leif

     
  • Leif Mortenson

    Leif Mortenson - 2007-03-17
    • labels: --> Misc
    • assigned_to: nobody --> mortenson
     
  • Nobody/Anonymous

    Logged In: NO

    Hi Leif,

    Is it not better then to use separate threads for pinging the JVM and for logging its output to disk? In my opinion, pinging the JVM is VERY important, while logging its output to the disk is not so important. If separate threads were used, the JVM would have been pinged although the disk froze up, and then the JVM would not have restarted itself.
    Logging can always be kept in memory for some time, and be written to the disk later.

    A hardware failure in the RAID controller probably is the cause of the freezing of the disks. It has operated well for two consequtive years now, and this is the first error on it. No further clues about the cause.

    Greetings,
    Johan
    wf_johan@yahoo.com

     

Log in to post a comment.