You can subscribe to this list here.
| 2001 |
Jan
|
Feb
(1) |
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
|
Sep
|
Oct
|
Nov
|
Dec
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2002 |
Jan
(1) |
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
|
Nov
(1) |
Dec
|
| 2003 |
Jan
|
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
(83) |
Nov
(57) |
Dec
(111) |
| 2004 |
Jan
(38) |
Feb
(121) |
Mar
(107) |
Apr
(241) |
May
(102) |
Jun
(190) |
Jul
(239) |
Aug
(158) |
Sep
(184) |
Oct
(193) |
Nov
(47) |
Dec
(68) |
| 2005 |
Jan
(190) |
Feb
(105) |
Mar
(99) |
Apr
(65) |
May
(92) |
Jun
(250) |
Jul
(197) |
Aug
(128) |
Sep
(101) |
Oct
(183) |
Nov
(186) |
Dec
(42) |
| 2006 |
Jan
(102) |
Feb
(122) |
Mar
(154) |
Apr
(196) |
May
(181) |
Jun
(281) |
Jul
(310) |
Aug
(198) |
Sep
(145) |
Oct
(188) |
Nov
(134) |
Dec
(90) |
| 2007 |
Jan
(134) |
Feb
(181) |
Mar
(157) |
Apr
(57) |
May
(81) |
Jun
(204) |
Jul
(60) |
Aug
(37) |
Sep
(17) |
Oct
(90) |
Nov
(122) |
Dec
(72) |
| 2008 |
Jan
(130) |
Feb
(108) |
Mar
(160) |
Apr
(38) |
May
(83) |
Jun
(42) |
Jul
(75) |
Aug
(16) |
Sep
(71) |
Oct
(57) |
Nov
(59) |
Dec
(152) |
| 2009 |
Jan
(73) |
Feb
(213) |
Mar
(67) |
Apr
(40) |
May
(46) |
Jun
(82) |
Jul
(73) |
Aug
(57) |
Sep
(108) |
Oct
(36) |
Nov
(153) |
Dec
(77) |
| 2010 |
Jan
(42) |
Feb
(171) |
Mar
(150) |
Apr
(6) |
May
(22) |
Jun
(34) |
Jul
(31) |
Aug
(38) |
Sep
(32) |
Oct
(59) |
Nov
(13) |
Dec
(62) |
| 2011 |
Jan
(114) |
Feb
(139) |
Mar
(126) |
Apr
(51) |
May
(53) |
Jun
(29) |
Jul
(41) |
Aug
(29) |
Sep
(35) |
Oct
(87) |
Nov
(42) |
Dec
(20) |
| 2012 |
Jan
(111) |
Feb
(66) |
Mar
(35) |
Apr
(59) |
May
(71) |
Jun
(32) |
Jul
(11) |
Aug
(48) |
Sep
(60) |
Oct
(87) |
Nov
(16) |
Dec
(38) |
| 2013 |
Jan
(5) |
Feb
(19) |
Mar
(41) |
Apr
(47) |
May
(14) |
Jun
(32) |
Jul
(18) |
Aug
(68) |
Sep
(9) |
Oct
(42) |
Nov
(12) |
Dec
(10) |
| 2014 |
Jan
(14) |
Feb
(139) |
Mar
(137) |
Apr
(66) |
May
(72) |
Jun
(142) |
Jul
(70) |
Aug
(31) |
Sep
(39) |
Oct
(98) |
Nov
(133) |
Dec
(44) |
| 2015 |
Jan
(70) |
Feb
(27) |
Mar
(36) |
Apr
(11) |
May
(15) |
Jun
(70) |
Jul
(30) |
Aug
(63) |
Sep
(18) |
Oct
(15) |
Nov
(42) |
Dec
(29) |
| 2016 |
Jan
(37) |
Feb
(48) |
Mar
(59) |
Apr
(28) |
May
(30) |
Jun
(43) |
Jul
(47) |
Aug
(14) |
Sep
(21) |
Oct
(26) |
Nov
(10) |
Dec
(2) |
| 2017 |
Jan
(26) |
Feb
(27) |
Mar
(44) |
Apr
(11) |
May
(32) |
Jun
(28) |
Jul
(75) |
Aug
(45) |
Sep
(35) |
Oct
(285) |
Nov
(99) |
Dec
(16) |
| 2018 |
Jan
(8) |
Feb
(8) |
Mar
(42) |
Apr
(35) |
May
(23) |
Jun
(12) |
Jul
(16) |
Aug
(11) |
Sep
(8) |
Oct
(16) |
Nov
(5) |
Dec
(8) |
| 2019 |
Jan
(9) |
Feb
(28) |
Mar
(4) |
Apr
(10) |
May
(7) |
Jun
(4) |
Jul
(4) |
Aug
|
Sep
(4) |
Oct
|
Nov
(23) |
Dec
(3) |
| 2020 |
Jan
(19) |
Feb
(3) |
Mar
(22) |
Apr
(17) |
May
(10) |
Jun
(69) |
Jul
(18) |
Aug
(23) |
Sep
(25) |
Oct
(11) |
Nov
(20) |
Dec
(9) |
| 2021 |
Jan
(1) |
Feb
(7) |
Mar
(9) |
Apr
|
May
(1) |
Jun
(8) |
Jul
(6) |
Aug
(8) |
Sep
(7) |
Oct
|
Nov
(2) |
Dec
(23) |
| 2022 |
Jan
(23) |
Feb
(9) |
Mar
(9) |
Apr
|
May
(8) |
Jun
(1) |
Jul
(6) |
Aug
(8) |
Sep
(30) |
Oct
(5) |
Nov
(4) |
Dec
(6) |
| 2023 |
Jan
(2) |
Feb
(5) |
Mar
(7) |
Apr
(3) |
May
(8) |
Jun
(45) |
Jul
(8) |
Aug
|
Sep
(2) |
Oct
(14) |
Nov
(7) |
Dec
(2) |
| 2024 |
Jan
(4) |
Feb
(4) |
Mar
|
Apr
(7) |
May
(2) |
Jun
(1) |
Jul
|
Aug
(5) |
Sep
|
Oct
|
Nov
(4) |
Dec
(14) |
| 2025 |
Jan
(22) |
Feb
(6) |
Mar
(5) |
Apr
(14) |
May
(6) |
Jun
(11) |
Jul
(19) |
Aug
|
Sep
(17) |
Oct
(1) |
Nov
(2) |
Dec
(18) |
| 2026 |
Jan
|
Feb
|
Mar
(5) |
Apr
|
May
(2) |
Jun
(1) |
Jul
(6) |
Aug
(1) |
Sep
|
Oct
|
Nov
|
Dec
|
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-04 17:14:51
|
On Saturday 04 June 2005 05:00 am, Jonathan Thornburg wrote: > On Thu, 2 Jun 2005, Ethan Merritt wrote: > > On Wednesday 25 May 2005 05:03 am, Lars Hecking wrote: > >> may I ask to advance the #define'd compile-time parameter > >> MAX_NUM_VAR > >> in src/syscfg.h from 5 to something like 20 or so? :-) > > How about changing these arrays to hold dynamically > > allocated pointers instead. [[...]] > > Is it worth complexifying the code to save 1K bytes of memory? No. But that isn't the motivation. The idea is to make parameter allocation dynamic, so that we don't have a hard-coded maximum number. I was just pointing out that switching to dynamic allocation would save space as a side benefit. > Does current gnuplot still run on MS-DOS 16-bit systems? There was some debate about this during the runup to releasing version 4. As currently organized, the code is too big for 16-bit DOS. This could be fixed without too much trouble by splitting out large chunks of driver-specific data into separate files. However, we took a poll at the time, and turned up not a single user interested in a 16-bit version. So we didn't bother. As it happens, I favor breaking out at least the PostScript prolog code into a separate file anyhow so that it can be customized to fit the needs of a specific site without recompiling gnuplot. But I am told that introduces installation problems under MSWin because there is no convention for where applications should place or look for associated data files. -- Ethan A Merritt Biomolecular Structure Center University of Washington 98195-7742 |
|
From: Jonathan T. <jt...@ae...> - 2005-06-04 12:00:40
|
On Thu, 2 Jun 2005, Ethan Merritt wrote:
> On Wednesday 25 May 2005 05:03 am, Lars Hecking wrote:
>>
>> may I ask to advance the #define'd compile-time parameter
>> MAX_NUM_VAR
>> in src/syscfg.h from 5 to something like 20 or so? :-)
>
> The arrays that use MAX_NUM_VAR are horribly wasteful,
> particularly if we bump it up to a larger number.
> e.g.: char c_dummy_var[MAX_NUM_VAR][MAX_ID_LEN+1];
> would be reserving 20*50 chars of storage to hold what
> is usually 2-3 instances of single-characters names like
> "u", "v" or "w".
>
> How about changing these arrays to hold dynamically
> allocated pointers instead. [[...]]
Is it worth complexifying the code to save 1K bytes of memory?
Does current gnuplot still run on MS-DOS 16-bit systems?
ciao,
--
-- Jonathan Thornburg <jt...@ae...>
Max-Planck-Institut fuer Gravitationsphysik (Albert-Einstein-Institut),
Golm, Germany, "Old Europe" http://www.aei.mpg.de/~jthorn/home.html
"Washing one's hands of the conflict between the powerful and the
powerless means to side with the powerful, not to be neutral."
-- quote by Freire / poster by Oxfam
|
|
From: Dimitrios A. <ji...@gm...> - 2005-06-03 21:21:03
|
> On Thursday 02 June 2005 03:53 pm, Dimitrios Apostolou wrote: >=20 >>So I submit to you a patch (against the v. 4.0 gnuplot) for the file=20 >>src/datafile.c as a proof of concept >=20 >=20 > Could you please re-do the patch using > diff -ur <oldfile> <newfile> >=20 I submit the patch for a third time, sorry for the spamming, but since=20 I=C2=B4m not subscribed to the list my emails wait for approval. This time I used the format you suggested: diff -ur <oldfile> <newfile> >>There are many things in my code that you 'll not probably like. >=20 >=20 > Your comments worry me. For instance: >=20 > < /* malloc the maximum we may use, it's ok in an overcommiting OS lik= e linux */ > < tmp_arr =3D gp_alloc(filesize * sizeof(float), "df_matrix"); >=20 > gnuplot core code must run on systems other than linux. > And even for linux your statement is not true. Many people doing > serious number crunching will not run linux in overcommit mode, because > it is too painful to see a computation which has already run for 3 days > be killed by the OOM killer just because someone has opened a web brows= er, > or in this case because they try to run gnuplot. As I said my code only serves as a proof of concept. In case you care to=20 use it this is one of the things that should change. Dimitris |
|
From: Xavier <xg...@lb...> - 2005-06-03 19:25:01
|
I read that the new version does not support this option, however I'd like to know if someone knows a work arouund to this problem. ~X |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-03 18:54:20
|
I have not been following this thread, so let me ask a few questions. You say there is a speed-up of 10X. What were the actual times as reported by the "time" command? Could you provide a profile analysis of the run time, so we could see where your CPU usage is being spent? Speeding things up from 2 seconds to 0.2 seconds, for example, is not very important. Or let me say that differently -- 2 seconds to read a file may be prohibitively slow if you are trying to use the mouse for interactive rotation, because currently the file is re-read at each mouse increment. But the proper fix for this, IMHO, is not to fiddle with the file reading code. Instead we should modify the replot command so that it re-uses the data previously read in if at all possible. On Thursday 02 June 2005 03:53 pm, Dimitrios Apostolou wrote: > > So I submit to you a patch (against the v. 4.0 gnuplot) for the file > src/datafile.c as a proof of concept Could you please re-do the patch using diff -ur <oldfile> <newfile> > There are many things in my code that you 'll not probably like. Your comments worry me. For instance: < /* malloc the maximum we may use, it's ok in an overcommiting OS like linux */ < tmp_arr = gp_alloc(filesize * sizeof(float), "df_matrix"); gnuplot core code must run on systems other than linux. And even for linux your statement is not true. Many people doing serious number crunching will not run linux in overcommit mode, because it is too painful to see a computation which has already run for 3 days be killed by the OOM killer just because someone has opened a web browser, or in this case because they try to run gnuplot. -- Ethan A Merritt merritt@u.washington.edu Biomolecular Structure Center Mailstop 357742 University of Washington, Seattle, WA 98195 |
|
From: Dimitrios A. <ji...@gm...> - 2005-06-03 17:56:37
|
Here is the unified patch. In the previous patch I also diffed the files in the wrong order. Dimitris |
|
From: Dimitrios A. <ji...@gm...> - 2005-06-03 17:48:23
|
Daniel J Sebald wrote: > Please *explain* why the patch is faster. Those listening will > understand. Also, when running diff be sure to use unified (-u) so that > it indicates what file the hunks come from. I don't know why it is faster. I just wrote a simple parser. I don't understand what more the old parser does. I just can see that it is much more complicated. My guess is that after so many years of development and after many additions that today we see but can't figure out, the code became a bit "bloated". Do what you think is better: optimize the current parser, rewrite a new one, or use mine as a base for improvement. One thing I know for sure: it shouldn't stay as it is. The patch I published is against the file: gnuplot-4.0.0/src/datafile.c Dimitris |
|
From: Daniel J S. <dan...@ie...> - 2005-06-03 16:37:35
|
Dimitrios Apostolou wrote: > So I submit to you a patch (against the v. 4.0 gnuplot) for the file > src/datafile.c as a proof of concept, that the current parser is slow > and can be improved. And you 'll see that this version is faster more > than 10 times. Please forgive any silly programming mistakes, I'm not > much experienced in C. Please *explain* why the patch is faster. Those listening will understand. Also, when running diff be sure to use unified (-u) so that it indicates what file the hunks come from. Dan |
|
From: Hans-Bernhard B. <br...@ph...> - 2005-06-03 10:13:39
|
Ga=EBl Varoquaux wrote:
> I am currently running Gnuplot on a filename with an output name with=
> several dot, eg : foo.bar :
> set output foo.bar.tex
My immediate answer is: don't do that. It's going to cause confusion,=20
and it's not necessary to do it that way, so don't do it that way.
> If I use the epslatex terminal it generates a
> eps file called foo.bar.eps, and a tex file in which there is the line =
:
>=20
> \includegraphics{foo.bar}%
>=20
> Now I don't know the graphicx package very well, but it seems that it=
> tries to load the file foo.bar and not the file foo.bar.eps, because it=
> is confused by the extra dot.=20
It's not confused, it follows its documented behaviour: if the filename=20
has an extension, use the name as-is. If not, try several extensions=20
from a (configurable) list of allowed ones, which includes .eps in a=20
plain LaTeX run, so {foo_bar} would find foo_bar.eps just fine.
> not compile. I suggest changing the epslatex driver so that it prints
> the line :
>=20
> \includegraphics{foo.bar.eps}%
>=20
> in the file to avoid this problem.
Absolutely no. This would kill one of the nicest features of the=20
epslatex driver: the fact that you just have to ps2pdf the eps part, and =
the LaTeX part will pass pdflatex flawlessly.
|
|
From: Dimitrios A. <ji...@gm...> - 2005-06-02 22:53:22
|
Hello. It would be nice if gnuplot parsed only the numbers it needed, but I understand this is not a priority. This is indeed a problem that can (should?) be corrected in the datafile. However, the second problem I noticed in gnuplot was the very slow parsing. I believe this is something that needs to be corrected. I tried to improve the parsing speed but I couldn't understand *many* things in the existing code. I 'm sure those are used somewhere but since I couldn't I understand them, I rewrote all the parser, the simplest way possible. So I submit to you a patch (against the v. 4.0 gnuplot) for the file src/datafile.c as a proof of concept, that the current parser is slow and can be improved. And you 'll see that this version is faster more than 10 times. Please forgive any silly programming mistakes, I'm not much experienced in C. There are many things in my code that you 'll not probably like. They were added to bypass several problems that the old codebase created. After all I only wrote this code as a proof of concept. However, if you think that this can replace the existing parser, tell me so, because there are a few things that need to change. Of course you may improve it as you wish. Dimitris |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-02 20:04:32
|
On Wednesday 25 May 2005 05:03 am, Lars Hecking wrote:
>
> may I ask to advance the #define'd compile-time parameter
> MAX_NUM_VAR
> in src/syscfg.h from 5 to something like 20 or so? :-)
The arrays that use MAX_NUM_VAR are horribly wasteful,
particularly if we bump it up to a larger number.
e.g.: char c_dummy_var[MAX_NUM_VAR][MAX_ID_LEN+1];
would be reserving 20*50 chars of storage to hold what
is usually 2-3 instances of single-characters names like
"u", "v" or "w".
How about changing these arrays to hold dynamically
allocated pointers instead. Then we would replace existing
assignments
strcpy(set_dummy_var[0], "x");
with calls to a routine
load_dummy_var(char *var, char *name)
{
free(var);
var = gp_strdup(name);
}
--
Ethan A Merritt merritt@u.washington.edu
Biomolecular Structure Center
Mailstop 357742
University of Washington, Seattle, WA 98195
|
|
From: V. <gae...@en...> - 2005-06-02 19:53:06
|
Hello,
I am currently running Gnuplot on a filename with an output name with
several dot, eg : foo.bar :
set output foo.bar.tex
If I use the epslatex terminal it generates a
eps file called foo.bar.eps, and a tex file in which there is the line :
\includegraphics{foo.bar}%
Now I don't know the graphicx package very well, but it seems that it
tries to load the file foo.bar and not the file foo.bar.eps, because it
is confused by the extra dot. Manually edit foo.bar.tex to change
"foo.bar" to "foo.bar.eps" works, but the out of the box tex file does
not compile. I suggest changing the epslatex driver so that it prints
the line :
\includegraphics{foo.bar.eps}%
in the file to avoid this problem.
Have I been clear ? (I don't really think so). Is there a workaround ?
Is there anything wrong with having in the tex file foo.bar.eps rather
than foo.bar.
Ga=EBl
PS : if you are interested in my case the name foo.bar is generated by a
script and is FORT.max.auto29 !
|
|
From: Daniel J S. <dan...@ie...> - 2005-05-30 17:11:50
|
Hans-Bernhard Broeker wrote: > Daniel J Sebald wrote: > >> Hans-Bernhard Broeker wrote: > > >>> Necessity is not the issue --- convenience of maintaining the overall >>> structure of a very ancient code base is. The code has always been >>> organized to read all data, then let the "every" filter decide which >>> points to actually use. > > >> Actually, from what I remember, in datafile.c the "everypoint" >> variable is used to toss out points. > > > Yes. But the question is: *when* does that happen: before sscanf()ing > the values, or afterwards? After a quick look into the current sources, > it would seem that at least for 'matrix' files, it reads the entire > thing before applying 'every'. That's probably the reason why 'every > 500:500' didn't achieve any speedup for Dimitrios. Oh yeah, I see now... and my memory is coming back. This is true in the case of binary data as well. I can't recall if I originated the concept or carried it over from previous code, but the idea was to bring in all the data and create a data format similar to what the normal routines read, but in memory. That could be changed I guess. (But would have to think of the ramifications.) The decimation could be done in the df_read_matrix() routine and then indicate to the main routine that "every" should be 1 instead of, say, 500. I wouldn't call that an urgent change, however, seeing as I have a number of things to do right now. Dan |
|
From: Hans-Bernhard B. <br...@ph...> - 2005-05-30 10:31:54
|
Daniel J Sebald wrote: > Hans-Bernhard Broeker wrote: >> Necessity is not the issue --- convenience of maintaining the overall >> structure of a very ancient code base is. The code has always been >> organized to read all data, then let the "every" filter decide which >> points to actually use. > Actually, from what I remember, in datafile.c the "everypoint" variable > is used to toss out points. Yes. But the question is: *when* does that happen: before sscanf()ing the values, or afterwards? After a quick look into the current sources, it would seem that at least for 'matrix' files, it reads the entire thing before applying 'every'. That's probably the reason why 'every 500:500' didn't achieve any speedup for Dimitrios. |
|
From: Daniel J S. <dan...@ie...> - 2005-05-29 21:03:58
|
Hans-Bernhard Broeker wrote: >> >> IMHO the more points we have the better looks the map or the surface >> mesh we plot. > > > That assumption is fatally flawed --- as soon as you have as many input > points than the output medium has pixels, adding more is guaranteed to > make the plot not better, but will actually render it increasingly > unreadable. 6000x6000 is well beyond that point: on a screen, you'll be > trying to display at least 36 data points in every pixel of your plot > --- that's not adding quality, that's adding confusion. Hans is right, Dimitris. There is also the very important issue of properly processing the data before discarding anything. I'm not a fan of the "every" qualifier, but I put it into the binary input functionality because it already existed in the ascii input methods. For images (or mesh plots, whatever) there is the concept of spatial frequency, i.e., how quickly intensity or color changes from pixel to pixel. If one simply disregards that and tosses out every other pixel, or only keeps every 500th pixel, etc., there is the possibility of experiencing aliasing. This means that some high spatial frequency part of the image could be aliased to a lower frequency, something that wasn't there previously. The effect can be very bad. The proper processing is to first low-pass filter the image with some form of 2D kernel, THEN toss out pixels. In all likelihood, printers must do this step if you send it a 6000 x 6000 image. Same would hold for a proper PostScript screen viewer. However, if you know that the data in the file you are going to decimate is of sufficiently low frequency then sure, you can keep every Nth pixel without harm. (But if that is the case then why store such a high resolution image? Anyway...) If not, you should process the thing in Matlab first and create a downsampled data set. As someone in this discussion pointed out. Once we have gnuplot doing all the low-pass filtering and so on suddenly gnuplot is more than just a plotting program. > >>> That's because gnuplot parses all data points, regardless of whether > > >> Is it really necessary? Why not parse only the needed points? > > > Necessity is not the issue --- convenience of maintaining the overall > structure of a very ancient code base is. The code has always been > organized to read all data, then let the "every" filter decide which > points to actually use. Changing that would be difficult, to say the > least. Features like the fact that 'using' can be applied even to > 'matrix' data may well rely on such details. Actually, from what I remember, in datafile.c the "everypoint" variable is used to toss out points. Not completely sure on that, but that would mean that gnuplot does not internally store those points not used as a consequence of "every". I believe "binary" works the same way with "every". An hour does seem ridiculously long even for a slower machine. If you are running out of memory linux will slow to a crawl because it is always swapping memory on and off the hard drive. Dan |
|
From: KITA T. <t-...@cc...> - 2005-05-29 15:12:26
|
From: Ga=EBl Varoquaux <gae...@en...> Subject: Povray terminal driver Date: Thu, 12 May 2005 02:22:32 +0200 > http://www.eleves.ens.fr/home/varoquau/gnuplot/ > = > Here is the first draft of the povray terminal driver. It is fully = 2D, > wright now. A lot of work remains to be done but I am going away for = two > weeks; Among the big piece of work to do is the font handling and > enhanced text bit. I have added RGb color support a few days ago have= > have ever since lost color. # Long time since I put the povrml patch. # Spring is the busiest season for university teachers in Japan. It is a good starting point for me and for the project to fit the 3D terminals to the gnuplot core, I believe. # Thank you for leaving my name in the code, though my code does not se= ems to # be of much help to you. To all who joined the discussion started from my patch posted 2 month a= go: Thank you very much. I read all the comments posted but I have no time to understand and to think how to reply to each referring the codes so far. I would like to continue revise my dirty patch considering all of your comments and suggestions. I also expect a new 3D terminal API/framework in the gnuplot core. Excuse me for the impolite response like this, but maybe better than no= thing. # The OpenGL terminal seems another interesting one, though I have not = yet # tried to build it. -- = KITA Toshihiro http://t-kita.net/ PGP-Key: http://t-kita.net/pubkey.asc fingerprint : CBFF 6A61 5990 10F5 B4B6 D2E8 279A 7063 CF8B 6339 |
|
From: Hans-Bernhard B. <br...@ph...> - 2005-05-29 14:06:13
|
Dimitrios Apostolou wrote: > Hans-Bernhard Broeker wrote: >> Dimitrios Apostolou wrote: > I like gnuplot so I tried it. For this kind of data I really like the > "map" capability of gnuplot. It would be interesting if the > "convenience" features worked faster. Interesting, yes. But we're already stretched somewhat thin on man-power as it is. Concerning ourselves with the efficiency of side aspects that other tools are already a lot better at than gnuplot can possibly be, would be a waste of effort. Note that I'm not trying to tell you not to use gnuplot at all --- I'm trying to convey the message that you should use more tools than only gnuplot. gnuplot is good for plotting, but mediocre at mass data processing. So use better tools for that part of the job, then come back to gnuplot with the actualy plottable data. >> It's not *that* much more, actually. A double-precision variable >> takes 8 bytes, that's about three times as much as your ASCII data. >> Add the implied x and y variables missing in your matrix file and >> you're at >> 6000*6000*3*8 Bytes = 864 MB of data. gnuplot will use even more than > > > Is *3 necessary for this kind of data (matrix)? For reasons of program structure and internal efficiency, all the various kinds of input data have to end up in the *same* data structure, regardless of whether they were topologically and geometrically very limited matrix data, generic grid-topology data, or a point cloud without any kind of structure. It's not strictly necessary to do it that way, but for most reasonable plots, this organization works well. >> that, and that's a problem. But the real problem here is that a >> 6000x6000 points data set is essentially unplottable --- no output >> device you're likely to be using has enough resolution to display all >> those points in a readable way. > > > IMHO the more points we have the better looks the map or the surface > mesh we plot. That assumption is fatally flawed --- as soon as you have as many input points than the output medium has pixels, adding more is guaranteed to make the plot not better, but will actually render it increasingly unreadable. 6000x6000 is well beyond that point: on a screen, you'll be trying to display at least 36 data points in every pixel of your plot --- that's not adding quality, that's adding confusion. >> That's because gnuplot parses all data points, regardless of whether > Is it really necessary? Why not parse only the needed points? Necessity is not the issue --- convenience of maintaining the overall structure of a very ancient code base is. The code has always been organized to read all data, then let the "every" filter decide which points to actually use. Changing that would be difficult, to say the least. Features like the fact that 'using' can be applied even to 'matrix' data may well rely on such details. >> they'll be used or not --- and, like it or not, scanning ASCII >> representations of (presumably) floating-point numbers is *slow*. > I know of the overhead "parsing" implies, however I know that an 800 Mhz > CPU ought to do it much faster. Probably. I just ran a little experiment, and found that 36000000 double-precision numbers could be scanf()ed in about 80 seconds process CPU time, on a 650 MHz PIII. It may be worthwile to profile your gnuplot in (a smaller version of) this case, to see where the time is actually spent. Possibly, it's pure memory access time --- at this kind of size, memory bandwidth becomes a serious bottleneck, too. > Don't you agree that gnuplot's > implementation is highly inefficient on this? ASCII datafiles are inefficient by design, but at the same time, this inefficiency allows them to be understood by humans, and keeps them 100% portable across machine architectures. Two sides of the same medal. |
|
From: V. <gae...@en...> - 2005-05-27 16:51:00
|
Hello, Have you tried using octave to preprocess the matrix before sending it to Gnuplot via a buffer file or throught the Octave/Gnuplot interface. Gnuplot is not a math program, it is a plotting program. If you want to process huge amount of data use a math program (or for simple matrice rechaping I found out awk is quite convenient). -- Ga=EBl |
|
From: Dimitrios A. <ji...@gm...> - 2005-05-27 12:16:57
|
Thank you all for your answers. Hans-Bernhard Broeker wrote: > Dimitrios Apostolou wrote: > >> I just notice that gnuplot (4.0) is extremely slow when dealing with >> big files. In particular I execute the command: >> >> splot 'matrix.asc' matrix every 500:500 >> >> where matrix.asc is an 130MB file containing a 6000x6000 matrix. > > > That's exceptionally little data per point. 130MB/(6000x6000) leaves > only 3 bytes/point, i.e. 2 decimal digits. In case you care, what I try to do is a quick hack to plot SRTM data. If you want to reproduce my exact steps do the following: - download a file from ftp://srtm.csi.cgiar.org/SRTM_Data_ArcAscii/ and unzip it - sed -n '/^[0-9\-].*/p' thefile.asc | sed 's/-9999/0/g' > matrix.asc - gnuplot - splot 'matrix.asc' matrix every 500:500 >> What I don't like is that altough the points to plot are about 12x12 >> the processing takes about half an hour. > > > So don't use gnuplot for it. 'using', 'every' are convenience features > for quick and easy on-the fly data selection and manipulation, not the > ultimate data processing tool. For that, use awk, perl, a spreadsheet, > or whatever floats your boat. I like gnuplot so I tried it. For this kind of data I really like the "map" capability of gnuplot. It would be interesting if the "convenience" features worked faster. >> If I don't specify "every 500:500" the gnuplot process uses more than >> 1GB of memory (after much time) and gets killed by the OS. So a second >> point is that it uses more memory than necessary. > > > It's not *that* much more, actually. A double-precision variable takes > 8 bytes, that's about three times as much as your ASCII data. Add the > implied x and y variables missing in your matrix file and you're at > 6000*6000*3*8 Bytes = 864 MB of data. gnuplot will use even more than Is *3 necessary for this kind of data (matrix)? > that, and that's a problem. But the real problem here is that a > 6000x6000 points data set is essentially unplottable --- no output > device you're likely to be using has enough resolution to display all > those points in a readable way. IMHO the more points we have the better looks the map or the surface mesh we plot. Of course I won't plot every point individually but as part of a surface. >> In both cases it is noteworthy that the hard disk is almost idle but >> the CPU at 100% all the time. What I mean is that the reading of the >> file is happening really slowly. > > > That's because gnuplot parses all data points, regardless of whether Is it really necessary? Why not parse only the needed points? > they'll be used or not --- and, like it or not, scanning ASCII > representations of (presumably) floating-point numbers is *slow*. I know of the overhead "parsing" implies, however I know that an 800 Mhz CPU ought to do it much faster. Don't you agree that gnuplot's implementation is highly inefficient on this? Please don't be offended by my comments. I think gnuplot is a very nice program and I only try to make it better. Of course sending a patch to you would be better but I'm not at all familiar with its code. > The problem is with the datafile, so that's where the solution has to > be. Use external tools to reduce it to a manageable size. I will do it, thanks. Or perhaps I will try to convert the datafile to binary format like someone else proposed. Thank you all for your answers, Dimitris |
|
From: Hans-Bernhard B. <br...@ph...> - 2005-05-27 10:53:16
|
Dimitrios Apostolou wrote: > I just notice that gnuplot (4.0) is extremely slow when dealing with big > files. In particular I execute the command: > > splot 'matrix.asc' matrix every 500:500 > > where matrix.asc is an 130MB file containing a 6000x6000 matrix. That's exceptionally little data per point. 130MB/(6000x6000) leaves only 3 bytes/point, i.e. 2 decimal digits. > What I don't like is that altough the points to plot are about 12x12 > the processing takes about half an hour. So don't use gnuplot for it. 'using', 'every' are convenience features for quick and easy on-the fly data selection and manipulation, not the ultimate data processing tool. For that, use awk, perl, a spreadsheet, or whatever floats your boat. > If I don't specify "every 500:500" the gnuplot process uses more than > 1GB of memory (after much time) and gets killed by the OS. So a second > point is that it uses more memory than necessary. It's not *that* much more, actually. A double-precision variable takes 8 bytes, that's about three times as much as your ASCII data. Add the implied x and y variables missing in your matrix file and you're at 6000*6000*3*8 Bytes = 864 MB of data. gnuplot will use even more than that, and that's a problem. But the real problem here is that a 6000x6000 points data set is essentially unplottable --- no output device you're likely to be using has enough resolution to display all those points in a readable way. > In both cases it is noteworthy that the hard disk is almost idle but the > CPU at 100% all the time. What I mean is that the reading of the file is > happening really slowly. That's because gnuplot parses all data points, regardless of whether they'll be used or not --- and, like it or not, scanning ASCII representations of (presumably) floating-point numbers is *slow*. The problem is with the datafile, so that's where the solution has to be. Use external tools to reduce it to a manageable size. |
|
From: Petr M. <mi...@ph...> - 2005-05-27 10:35:10
|
> I just notice that gnuplot (4.0) is extremely slow when dealing with big > where matrix.asc is an 130MB file containing a 6000x6000 matrix. What I I think writing such a huge file is very slow as well. Why don't you use gnuplot binary format? See 'help binary'. You can decreses memory consumption by setting in syscfg.h COORDVAL_FLOAT from double to float. Finally, consider using the development version 4.1. It supports images, binary image files etc. --- PM |
|
From: Petr M. <mi...@ph...> - 2005-05-27 09:30:22
|
> http://www.eleves.ens.fr/home/varoquau/gnuplot/ > > Have a look at it, enjoy, tell me how to improve it or finish it... I > continue working on it only in two weeks times. Looks interesting. Please add support for 'set palette maxcolors". --- PM |
|
From: Dimitrios A. <ji...@gm...> - 2005-05-26 22:19:13
|
Hello list, I just notice that gnuplot (4.0) is extremely slow when dealing with big files. In particular I execute the command: splot 'matrix.asc' matrix every 500:500 where matrix.asc is an 130MB file containing a 6000x6000 matrix. What I don't like is that altough the points to plot are about 12x12 the processing takes about half an hour. If I don't specify "every 500:500" the gnuplot process uses more than 1GB of memory (after much time) and gets killed by the OS. So a second point is that it uses more memory than necessary. Actually the memory consumed is enormous considering that the points are "only" 36.000.000. In both cases it is noteworthy that the hard disk is almost idle but the CPU at 100% all the time. What I mean is that the reading of the file is happening really slowly. Is this a bug? Or am I doing something wrong? Is there a workaround to speed things up? Thanks in advance, Dimitris P.S. Please CC replies directly to me since I'm not subscribed to the list |
|
From: Petr M. <mi...@ph...> - 2005-05-26 14:14:59
|
> > Consequently, I propose that gnuplot always reads .Xdefaults resources in > > C-locale. > > I fail to understand the problem. > It's your .Xdefaults file - you can put anything you like in it. > Why limit gnuplot by telling it to ingore the user's choice of > environment? I just wanted that gnuplot entries in .Xdefault are not subject to locale, but always with decimal point. Patch by Hans-Bernhard fixed this. > I think what you want is > setenv LC_CTYPE cs_CZ.UTF-8,en_US.UTF-8 Interesting to know that you can have two options. --- PM |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-05-25 16:27:42
|
On Wednesday 25 May 2005 12:27 am, Petr Mikulik wrote: > Consequently, I propose that gnuplot always reads .Xdefaults resources in > C-locale. I fail to understand the problem. It's your .Xdefaults file - you can put anything you like in it. Why limit gnuplot by telling it to ingore the user's choice of environment? > PS: I came to this issue because on my system there is by default > LANG=cs_CZ.UTF-8 > but I run Midnight Commander as > LANG=en_US.UTF-8 mc_lastdir > so that its menus are always in English. That's not funny if hotkeys change. Aha. So I think the underlying issue is that you are using obsolete syntax for setting the locale. For a long time now the locale settings have been split into separate components, and the LANG variable is used only as a backward-compatibility fall-back. I think what you want is unsetenv LANG setenv LC_CTYPE cs_CZ.UTF-8,en_US.UTF-8 setenv LC_MESSAGES en_US.UTF-8 setenv LC_NUMERIC C Individual programs may be buggy, of course, but if they are correctly calling setlocale() then this should result in using CZ character encodings if available (with fall-back to en_US), menus and other system messages in English, and C-language formats for numbers. You can also set LC_TIME to choose a prefered time/date format. > Gnuplot behaves differently if I run it from plain Konsole or from mc, > i.e. its .Xdefault resources are not portable. It should be consistent if you set LC_NUMERIC. If you find out otherwise, then that definitely is a bug, and we should try to fix it. I am not an expert on setting locales. I had to figure this all out the hard way, after installing my current generation of machines to support ja_JP.UTF-8. I assure you it's at least as disconcerting to have your menus pop up in Japanese as it is to have them appear in Czech. But dual-language support works nicely once you have the various locale settings sorted out. I love the current versions of Mandrake (now Mandriva), where the KDE desktop can be installed with SCIM multi-language support. This is really great because you can change the effective locale for individual programs (while they are running!) with a hot keys or a menu command. (caveat: the programs have to be build with appropriate X-input support, but that is true for the ones in the Mandriva distro). -- Ethan A Merritt merritt@u.washington.edu Biomolecular Structure Center Mailstop 357742 University of Washington, Seattle, WA 98195 |