Menu

#212 gLite submission of 100 jobs

3.6.1
open
imi
None
9 - High
2014-02-27
2014-02-26
Mahdi
No

Hi

I submitted a workflow with a generator that generates around 100 gLite jobs. I submitted this workflow with different split factors with a total of less than 200 jobs in parallel. It looks like gUSE is pretty overloaded now, and the problem is if DCI-Bridge fails (once?) to get the status of a job, it gives up and considers the job as ERROR. So far almost half of the jobs have failed with this reason, and half of the rest have finished, and obviously the rest are still running.

Furthermore, I see job items in the error list with messages like "upload error" or "20000 (get status failed)" in the "Resource" column (see attachment). Clicking on any buttons (Std. err or others) results in "Information not available !". In the attachment, you can see that there are 10 jobs in the list of errors, but the number shown in the back window for "error" is 1. Later this number is corrected as the number of failed and finished jobs grew. The above experience is in gUSE 3.5.8.

On a VM running gUSE 3.6.1, I submitted the exact same workflow again. This time I submitted twice, each time generating 100 jobs (total of 200 jobs). After one hour, the 200 jobs are still in "submitted" state. The VM, on the other hand, has become very unresponsive, making it impossible to check the logs from command line.

As a final remark, submitting a few hundred jobs is not outstanding at all, and we (at AMC) have an end-user who is planning to submit such an experiment (with around 100 jobs) in very near future, using our gUSE-based science gateway. I hope in next releases this will be more reliable, but I also appreciate any recommendation on how to deal with this problem in the current version.

Regards
Mahdi (AMC)

1 Attachments

Discussion

  • Zoltán Farkas

    Zoltán Farkas - 2014-02-27
    • assigned_to: imi
     
  • Gabor Hermann

    Gabor Hermann - 2014-02-27

    Dear Mahdi
    I assume that first of all the improper portal installation may be responsible for the experienced bad performance:
    As a thumb of rule the distributed portal installation and the proper queue number of the given DCI may dramatically improve the throughput.

    I have repeated your experiment in a simple test portal of mine based of the following conditions:

    • Host of the WS-PGRADE gUSE infrastructure : 4 –not too strong – processors, 2 GB RAM in a virtual machine of our Open Nebula Cloud.

    • Single machine portal installation.

    • Number of input queues of the gLITE submitter within the DCI-Bridge: 5

    I have experienced no blocking of the UIF response and the development of the job states seemed to be normal.
    I have executed several experiments where I have sent parallel 100 jobs instances into the seegrid VO.
    The jobs run for a predefined time and during that they perform floating point operations.

    Results:

    Result of 60 sec runs: The whole workflow terminated in 720 sec performing together 1.443661E+12 floating operations
    Result of 600 sec runs: The whole workflow terminated in 7317 sec performing together 1.362242E+13 floating operations

     

Log in to post a comment.