I originally had this problem with 3.0.2, and patched
our local copy. After upgrading to 3.1.2 I still had
the problem, but patching looked more difficult because
of changes to the signal handling infrastructure...
The wrapper is producing zombie processes, and the way
we fixed it was by installing a signal handler for
SIGCHLD signals.
Here was my explanation to the engineering managers in
our local bugzilla:
SIGCHLD is sent to the process by the kernel whenever
one of the processes' children is killed. In the
handler for SIGCHLD signals, which is just a
function, you need to call wait(NULL), to wait until
the process is officially
dead, so that the child's exit code is read by
something (see previos comment).
That's it! When these two small changes are made the
code no longer produces
zombie processes. See attachment on next message for
code ..
/**
* Handle child death
*/
void handleChildDeath(int sig_num) {
signal(SIGCHLD, handleChildDeath);
log_printf(WRAPPER_SOURCE_WRAPPER, LEVEL_STATUS,
"Received SIGCHLD, calling
wait().");
wait(NULL);
log_printf(WRAPPER_SOURCE_WRAPPER, LEVEL_STATUS,
"wait() returned, zombie
should be gone.");
}
/**
* Execute initialization code to get the wrapper set up.
*/
int wrapperInitialize() {
int retval = 0;
/* Set handlers for signals */
if (signal(SIGINT, handleInterrupt) == SIG_ERR ||
signal(SIGQUIT, handleQuit) == SIG_ERR ||
signal(SIGCHLD, handleChildDeath) == SIG_ERR ||
signal(SIGTERM, handleTermination) == SIG_ERR) {
retval = -1;
}
return retval;
}
Please let me know if you guys are willing to merge in
a fix to the primary CVS tree. Thanks.
Traun Leyden <tleyden@atypon.com>
Logged In: YES
user_id=228081
Traun,
Thanks for the patch. I had not realized how that worked.
I implemented your patch with a few modifications to fit the
latest version. It will be in the 3.2.0 release.
Cheers,
Leif