You can subscribe to this list here.
| 2001 |
Jan
|
Feb
(1) |
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
|
Sep
|
Oct
|
Nov
|
Dec
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2002 |
Jan
(1) |
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
|
Nov
(1) |
Dec
|
| 2003 |
Jan
|
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
(83) |
Nov
(57) |
Dec
(111) |
| 2004 |
Jan
(38) |
Feb
(121) |
Mar
(107) |
Apr
(241) |
May
(102) |
Jun
(190) |
Jul
(239) |
Aug
(158) |
Sep
(184) |
Oct
(193) |
Nov
(47) |
Dec
(68) |
| 2005 |
Jan
(190) |
Feb
(105) |
Mar
(99) |
Apr
(65) |
May
(92) |
Jun
(250) |
Jul
(197) |
Aug
(128) |
Sep
(101) |
Oct
(183) |
Nov
(186) |
Dec
(42) |
| 2006 |
Jan
(102) |
Feb
(122) |
Mar
(154) |
Apr
(196) |
May
(181) |
Jun
(281) |
Jul
(310) |
Aug
(198) |
Sep
(145) |
Oct
(188) |
Nov
(134) |
Dec
(90) |
| 2007 |
Jan
(134) |
Feb
(181) |
Mar
(157) |
Apr
(57) |
May
(81) |
Jun
(204) |
Jul
(60) |
Aug
(37) |
Sep
(17) |
Oct
(90) |
Nov
(122) |
Dec
(72) |
| 2008 |
Jan
(130) |
Feb
(108) |
Mar
(160) |
Apr
(38) |
May
(83) |
Jun
(42) |
Jul
(75) |
Aug
(16) |
Sep
(71) |
Oct
(57) |
Nov
(59) |
Dec
(152) |
| 2009 |
Jan
(73) |
Feb
(213) |
Mar
(67) |
Apr
(40) |
May
(46) |
Jun
(82) |
Jul
(73) |
Aug
(57) |
Sep
(108) |
Oct
(36) |
Nov
(153) |
Dec
(77) |
| 2010 |
Jan
(42) |
Feb
(171) |
Mar
(150) |
Apr
(6) |
May
(22) |
Jun
(34) |
Jul
(31) |
Aug
(38) |
Sep
(32) |
Oct
(59) |
Nov
(13) |
Dec
(62) |
| 2011 |
Jan
(114) |
Feb
(139) |
Mar
(126) |
Apr
(51) |
May
(53) |
Jun
(29) |
Jul
(41) |
Aug
(29) |
Sep
(35) |
Oct
(87) |
Nov
(42) |
Dec
(20) |
| 2012 |
Jan
(111) |
Feb
(66) |
Mar
(35) |
Apr
(59) |
May
(71) |
Jun
(32) |
Jul
(11) |
Aug
(48) |
Sep
(60) |
Oct
(87) |
Nov
(16) |
Dec
(38) |
| 2013 |
Jan
(5) |
Feb
(19) |
Mar
(41) |
Apr
(47) |
May
(14) |
Jun
(32) |
Jul
(18) |
Aug
(68) |
Sep
(9) |
Oct
(42) |
Nov
(12) |
Dec
(10) |
| 2014 |
Jan
(14) |
Feb
(139) |
Mar
(137) |
Apr
(66) |
May
(72) |
Jun
(142) |
Jul
(70) |
Aug
(31) |
Sep
(39) |
Oct
(98) |
Nov
(133) |
Dec
(44) |
| 2015 |
Jan
(70) |
Feb
(27) |
Mar
(36) |
Apr
(11) |
May
(15) |
Jun
(70) |
Jul
(30) |
Aug
(63) |
Sep
(18) |
Oct
(15) |
Nov
(42) |
Dec
(29) |
| 2016 |
Jan
(37) |
Feb
(48) |
Mar
(59) |
Apr
(28) |
May
(30) |
Jun
(43) |
Jul
(47) |
Aug
(14) |
Sep
(21) |
Oct
(26) |
Nov
(10) |
Dec
(2) |
| 2017 |
Jan
(26) |
Feb
(27) |
Mar
(44) |
Apr
(11) |
May
(32) |
Jun
(28) |
Jul
(75) |
Aug
(45) |
Sep
(35) |
Oct
(285) |
Nov
(99) |
Dec
(16) |
| 2018 |
Jan
(8) |
Feb
(8) |
Mar
(42) |
Apr
(35) |
May
(23) |
Jun
(12) |
Jul
(16) |
Aug
(11) |
Sep
(8) |
Oct
(16) |
Nov
(5) |
Dec
(8) |
| 2019 |
Jan
(9) |
Feb
(28) |
Mar
(4) |
Apr
(10) |
May
(7) |
Jun
(4) |
Jul
(4) |
Aug
|
Sep
(4) |
Oct
|
Nov
(23) |
Dec
(3) |
| 2020 |
Jan
(19) |
Feb
(3) |
Mar
(22) |
Apr
(17) |
May
(10) |
Jun
(69) |
Jul
(18) |
Aug
(23) |
Sep
(25) |
Oct
(11) |
Nov
(20) |
Dec
(9) |
| 2021 |
Jan
(1) |
Feb
(7) |
Mar
(9) |
Apr
|
May
(1) |
Jun
(8) |
Jul
(6) |
Aug
(8) |
Sep
(7) |
Oct
|
Nov
(2) |
Dec
(23) |
| 2022 |
Jan
(23) |
Feb
(9) |
Mar
(9) |
Apr
|
May
(8) |
Jun
(1) |
Jul
(6) |
Aug
(8) |
Sep
(30) |
Oct
(5) |
Nov
(4) |
Dec
(6) |
| 2023 |
Jan
(2) |
Feb
(5) |
Mar
(7) |
Apr
(3) |
May
(8) |
Jun
(45) |
Jul
(8) |
Aug
|
Sep
(2) |
Oct
(14) |
Nov
(7) |
Dec
(2) |
| 2024 |
Jan
(4) |
Feb
(4) |
Mar
|
Apr
(7) |
May
(2) |
Jun
(1) |
Jul
|
Aug
(5) |
Sep
|
Oct
|
Nov
(4) |
Dec
(14) |
| 2025 |
Jan
(22) |
Feb
(6) |
Mar
(5) |
Apr
(14) |
May
(6) |
Jun
(11) |
Jul
(19) |
Aug
|
Sep
(17) |
Oct
(1) |
Nov
(2) |
Dec
(18) |
| 2026 |
Jan
|
Feb
|
Mar
(5) |
Apr
|
May
(2) |
Jun
(1) |
Jul
(6) |
Aug
(1) |
Sep
|
Oct
|
Nov
|
Dec
|
|
From: Dimitrios A. <ji...@gm...> - 2005-06-09 23:10:54
|
Ethan Merritt wrote: > Instead we should modify the replot command so that it re-uses > the data previously read in if at all possible. > I believe this would be a nice fix. Since replot is mostly used in interactive mode, this should not break many scripts. And perhaps a "reread" command would be nice for rereading the input file. > gnuplot core code must run on systems other than linux. > And even for linux your statement is not true. Many people doing > serious number crunching will not run linux in overcommit mode, because > it is too painful to see a computation which has already run for 3 days > be killed by the OOM killer just because someone has opened a web browser, > or in this case because they try to run gnuplot. I agree with you and don't like overcommit mode myself. How do you turn it off? Dimitris |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-09 21:13:16
|
On Thursday 09 June 2005 04:14 am, Robert Hart wrote: > > #ifdef USE_STRTOD > char *fin; > df_column[df_no_cols].datum=strtod(s,&fin); > used=s-fin; > count=used?1:0; > #else > count = sscanf(s, "%lf%n", &df_column[df_no_cols].datum, &used); > #endif Somewhat more complete patch attached, adding a TBOOLEAN to toggle fortran_float handling. I've run it through the usual test suite, but haven't yet tried deliberately to break it with oddball input data. -- Ethan A Merritt merritt@u.washington.edu Biomolecular Structure Center Mailstop 357742 University of Washington, Seattle, WA 98195 |
|
From: Daniel J S. <dan...@ie...> - 2005-06-09 21:11:45
|
Ethan Merritt wrote:
> On Thursday 09 June 2005 09:12 am, Robert Hart wrote:
>=20
>>On Thu, 2005-06-09 at 21:18 +0800, Hans-Bernhard Br=C3=B6ker wrote:
>>
>>>Well, here's an ancient comment right taken directly from datafile.c:
>>>
>>> /* cannot trust strtod - eg strtod("-",&p) */
>=20
>=20
> I cannot make any sense of that comment either. ANSI says, in regard
> to scanf with %f
> "The format of the token should be that expected by the function
> strtod for a floating-point number that uses decimal notation".
> So there doesn't seem to be much room for scanf and strtod to parse
> numbers differently.
Perhaps this was mentioned, but a conjecture on the slowness of scanf() m=
ight be=20
that the number of output arguments is unknown and requires special handl=
ing,=20
whereas strtod() has only one output argument.
Dan
|
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-09 20:54:58
|
On Thursday 09 June 2005 09:12 am, Robert Hart wrote:
> On Thu, 2005-06-09 at 21:18 +0800, Hans-Bernhard Br=C3=B6ker wrote:
> >=20
> > Well, here's an ancient comment right taken directly from datafile.c:
> >=20
> > /* cannot trust strtod - eg strtod("-",&p) */
I cannot make any sense of that comment either. ANSI says, in regard
to scanf with %f
"The format of the token should be that expected by the function
strtod for a floating-point number that uses decimal notation".
So there doesn't seem to be much room for scanf and strtod to parse
numbers differently.
=2D-=20
Ethan A Merritt merritt@u.washington.edu
Biomolecular Structure Center
Mailstop 357742
University of Washington, Seattle, WA 98195
|
|
From: Daniel J S. <dan...@ie...> - 2005-06-09 20:31:20
|
Interesting. Of course, I think the same situation exists with scanf() in the non-matrix ascii input. That is what brought about the "binary" option in the first place. Sending images from octave to gnuplot was sloooooow... Dan Ethan Merritt wrote: > Not the factor of 10X that Dimitrios Apostolou was hoping for, > but up to a factor of 5X depending on the platform. > > > linux/x86 > --------- > ./scan_timing 100000 "3.5D 1E07" > Test atod: 0.059186 > Test sscanf: 0.150265 > Test strtod: 0.060353 > > DU/alpha > --------- > ./scan_timing 100000 "3.5D 1E07" > Test atod: 0.100689 > Test sscanf: 0.588027 > Test strtod: 0.124872 > > linux/AMD64 > ----------- > ./scan_timing 100000 "3.5D 1E07" > Test atod: 0.033294 > Test sscanf: 0.076438 > Test strtod: 0.037122 > > irix/MIPS R4400 > ---------------- > ./scan_timing 100000 "3.5D 1E07" > Test atod: 0.352085 > Test sscanf: 1.012914 > Test strtod: 0.357065 > > SunOS 4.1.4 1 sun4m > --------------------------- > ./scan_timing 100000 "3.5D 1E07" > Test atod: 2.245737 > Test sscanf: 4.014354 > Test strtod: 2.706037 |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-09 20:15:36
|
On Thursday 09 June 2005 09:12 am, Robert Hart wrote: > > I have attached my test program. I can't find any input on my system > that gives different results between sscanf and strtod. Very nice. Thank you. So strtod is faster on all platforms I have tried, with the largest difference on DU/alpha and the smallest difference on SunOS. Not the factor of 10X that Dimitrios Apostolou was hoping for, but up to a factor of 5X depending on the platform. linux/x86 --------- ./scan_timing 100000 "3.5D 1E07" Test atod: 0.059186 Test sscanf: 0.150265 Test strtod: 0.060353 DU/alpha --------- ./scan_timing 100000 "3.5D 1E07" Test atod: 0.100689 Test sscanf: 0.588027 Test strtod: 0.124872 linux/AMD64 ----------- ./scan_timing 100000 "3.5D 1E07" Test atod: 0.033294 Test sscanf: 0.076438 Test strtod: 0.037122 irix/MIPS R4400 ---------------- ./scan_timing 100000 "3.5D 1E07" Test atod: 0.352085 Test sscanf: 1.012914 Test strtod: 0.357065 SunOS 4.1.4 1 sun4m --------------------------- ./scan_timing 100000 "3.5D 1E07" Test atod: 2.245737 Test sscanf: 4.014354 Test strtod: 2.706037 -- Ethan A Merritt merritt@u.washington.edu Biomolecular Structure Center Mailstop 357742 University of Washington, Seattle, WA 98195 |
|
From: Daniel J S. <dan...@ie...> - 2005-06-09 18:03:08
|
Hans-Bernhard Br=F6ker wrote: >> Neither should it be a memory problem (something that would really=20 >> slow down Dimitris' machine perahps). The block of memory it consumes= =20 >> will not be much bigger than is eventually stored in memory. >=20 >=20 > For the case of plot "datafile" matrix every 500:500, it might... OK. Point taken. That could probably be fixed by doing the decimation f= irst=20 and then changing the "every" internally to 1. >=20 >> But, what could be a problem--in light of the discussion about FORTRAN= =20 >> doubles--is that instead of "floats" maybe this routine should be=20 >> storing doubles. =20 >=20 >=20 > Yes, it should. More to the point, it should store in type coordval. > That's what it's defined for. OK. Who want's to change that? Or should we hold off until everything e= lse is=20 worked out? >=20 > > Certainly a plotting program doesn't need the >=20 >> resolution of doubles.=20 >=20 >=20 > Oh, but it does --- if only because some people want to be able to zoom= =20 > in somewhat deeply. 1/DBL_FLT is less than the number of pixels on the= =20 > screen, and only a zoom factor of 1000 away from introducing aliasing=20 > artefacts. You're right. I forgot about zooming. >=20 >> Unless one writes a more sophisticated parser. Given how much testing= =20 >> is already done here, that would almost be as good an option. (But=20 >> perhaps not preferable.) >=20 >=20 > A parser wouldn't do --- we'ld need a dynamic parser *generator*, i.e.=20 > an engine that, given a set of 'using', 'every', 'index' and whatnot,=20 > will generate a parser on-the-fly and run that on the code. [...] > Not really. This is a test for "the first <n> using specifiers are jus= t=20 > 1..<n>, with (<n> <=3D 5)". I don't think there's much of a way of=20 > expressing that any more efficiently. It might make sense to move this > test to a different place, though. I think you are saying the same thing here, and sort of what I was gettin= g at.=20 Couldn't these tests be done before looping through the code to figure ou= t <n>=20 and then test the loop counter against <n>. Not sure... >=20 >> NEXT ISSUE: >=20 > [...] >=20 >> /* HBB 20001221: avoid breaking parsing of time/date >> * strings like 01Dec2000 that would be caused by >> * overwriting the 'D' with an 'e'... */ >=20 >=20 >> Isn't this a bug? =20 [...] > No. This would kill both the scanf() format string allowed to be passe= d=20 > in, if that contained some random 'q' characters, and, perhaps more=20 > importantly, Ethan's datastrings stuff. Oy. Guess that is so. Dan |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-09 17:40:04
|
On Thursday 09 June 2005 12:26 am, Hans-Bernhard Br=F6ker wrote: > Ethan Merritt wrote: >=20 > There's not a lot you could do with strtod() that sscanf() couldn't do=20 > just as well. But slower, apparently. > >>sscanf() and its %n format are used for a reason. >=20 > > But %n is non-portable, as documented by the OSK comment=20 >=20 > Just because some silly implementation is buggy doesn't exactly mean %n=20 > is unportable. %n is as portable as you can ever hope to be: it's an=20 > ANSI requirement. Surely you know the quote about standards: "Standards are nice because there are so many of them...". =20 It turns out that different ANSI revisions describe the behavior of %n differently, so it is not portable. Here is an=20 excerpt from the current man page: "Probably it is wise not to make any assumptions on the effect of %n conversions on the return value." The current code assumes that=20 count =3D sscanf(s, "%lf%n", ...) will set count to 1 if successful, but on some systems it will instead set count to 2. I do not know if this is the specific problem with OSK. strtod() is more portable than %n, and allows more complete=20 error reporting than sscanf(). =20 > I dimly remember this being about not running sscanf() on all columns of= =20 > a multi-column data file if no extended using specs are in use. If you > have "using 1:2:3", you can simply ignore columns 4 to 1000. If you=20 > have "using 1:2:(column(some_function($3)), datafile.c has to convert=20 > all 1000 of them. That is an interesting point, and I now wonder if these optimization lines conflict with the new column-based syntax elements like xticlabels(). I will have to experiment with this. =2D-=20 Ethan A Merritt merritt@u.washington.edu Biomolecular Structure Center Mailstop 357742 University of Washington, Seattle, WA 98195 |
|
From: Dimitrios A. <ji...@gm...> - 2005-06-09 16:49:44
|
Petr Mikulik wrote: > Can you please update it for the current datafile.c from cvs on sourceforge? > (There are minor rejects.) Here you go. Warning: I haven't checked if it compiles and it still doesn't replace the old code. However, I think that the functions df_read_matrix_jimis() replaces are df_read_matrix(), df_tokenise(), df_gets(). The reason I use fgetc() instead of fgets() is that one line might be very large to store it all at once. Speed shouldn't be reduced at all since stdio does buffered I/O. Moreover, gnuplot is really memory inefficient IMHO. To read 36M numbers for example it needs: (36M * sizeof(float)) + (36M * 3 * sizeof(float)) The first parenthesis is freed after creating the second one but the fact is that all this space is needed at copying time. Wouldn't it be interesting if we stored numbers from the file directly to gnuplot's format (second parenthesis)? That would result in 25% less memory requirements. Finally, if you like the idea of rewriting the parser (me or somebody else), the matrix format has to be clarified. Which lines are considered comments: the ones starting immediately with "#"? are empty lines allowed? Fortran numbers? What is a valid delimiter (all whitespace)? Thanks, Dimitris |
|
From: Robert H. <en...@no...> - 2005-06-09 16:12:20
|
On Thu, 2005-06-09 at 21:18 +0800, Hans-Bernhard Br=F6ker wrote:
> Robert Hart wrote:
>=20
> > strtod() is indeed significantly faster than sscanf.=20
>
> And I still don't see why that should be the case...
sscanf takes an arbitrary format string as an argument, which it must
parse every time it is called, it is also very generic so, I imagine,
ends up doing a lot of extra work to support this. If you wanted to you
could look through the source for libc and figure it out.
> > Can anybody come up with any cases where these two methods give
> > different results/error handling?
>=20
> Well, here's an ancient comment right taken directly from datafile.c:
>=20
> /* cannot trust strtod - eg strtod("-",&p) */
I've seen this comment, and tried it out. In my tests this isn't an
issue. The man page says:
If no conversion is performed, zero is returned and the value of n=
ptr
is stored in the location referenced by endptr.
So if(nptr!=3Dendptr) you know it failed.
Am I missing something? Are there other libc's that behave differently?
I have attached my test program. I can't find any input on my system
that gives different results between sscanf and strtod.
> > strtod is ANSI C, so should be
> > portable.
>=20
> So should sscanf().
Well comments in the code say "%n" doesn't work on OSK. I don't even
know what OSK is, but if strtod works where sscanf doesn't then that's a
win.
Rob
--=20
Robert Hart <en...@no...>
University of Nottingham
This message has been checked for viruses but the contents of an attachment
may still contain software viruses, which could damage your computer system:
you are advised to perform your own checks. Email communications with the
University of Nottingham may be monitored as permitted by UK legislation.
|
|
From:
<br...@ph...> - 2005-06-09 13:50:08
|
Daniel J Sebald wrote:
> I've looked at the current code a bit. It's quite involved, isn't it?
Yes. But then, so is the variety of inputs gnuplot can handle.
> For Dimitris' sake, there seems to be some details he may have not
> accounted for and would be difficult to add to the patch. So, would it
> pay for him to continue with the patch he's supplied? Or does it look
> like the current code will have to be the starting point?
The latter. Its main value (and that of Robert's later works) is to
prove that it might be possible to do it faster. Next, we have to find out
a) if the faster code doesn't break any existing behaviour
b) if it does, if and how it's possible to have both the old code's
flexibility and the new one's speed.
> Neither should it be a memory problem (something that would really slow
> down Dimitris' machine perahps). The block of memory it consumes will
> not be much bigger than is eventually stored in memory.
For the case of plot "datafile" matrix every 500:500, it might...
> But, what could be a problem--in light of the discussion about FORTRAN
> doubles--is that instead of "floats" maybe this routine should be
> storing doubles.
Yes, it should. More to the point, it should store in type coordval.
That's what it's defined for.
> Certainly a plotting program doesn't need the
> resolution of doubles.
Oh, but it does --- if only because some people want to be able to zoom
in somewhat deeply. 1/DBL_FLT is less than the number of pixels on the
screen, and only a zoom factor of 1000 away from introducing aliasing
artefacts.
> Unless one writes a more sophisticated parser. Given how much testing
> is already done here, that would almost be as good an option. (But
> perhaps not preferable.)
A parser wouldn't do --- we'ld need a dynamic parser *generator*, i.e.
an engine that, given a set of 'using', 'every', 'index' and whatnot,
will generate a parser on-the-fly and run that on the code.
> There are three tests here that can be done *before* reading data
They are.
> and combined into a single test within tokenise:
Not sure that that would achieve enough to be worth doing.
> (fast_columns == 0)
This one is essentially the Corey Satten optimization.
> (df_no_use_specs == 0)
> df_no_use_specs > 5
> and somehow my intuition tells me the following could be done in a
> better way:
>
> || ((df_no_use_specs > 0)
> && (use_spec[0].column == dfncp1
> || (df_no_use_specs > 1
> && (use_spec[1].column == dfncp1
> || (df_no_use_specs > 2
> && (use_spec[2].column == dfncp1
> || (df_no_use_specs > 3
> && (use_spec[3].column == dfncp1
> || (df_no_use_specs > 4 && (use_spec[4].column
> == dfncp1 || )
Not really. This is a test for "the first <n> using specifiers are just
1..<n>, with (<n> <= 5)". I don't think there's much of a way of
expressing that any more efficiently. It might make sense to move this
test to a different place, though.
> NEXT ISSUE:
[...]
> /* HBB 20001221: avoid breaking parsing of time/date
> * strings like 01Dec2000 that would be caused by
> * overwriting the 'D' with an 'e'... */
> Isn't this a bug?
Not to my knowledge, no.
> If so, why not?
Why should it be one?
> Certainly the date string 01Dec2000
> should not be interpretted as a double.
No. But it has to survive the attempted interpretation as a double,
unchanged. I.e. at this point of execution, the parsing engine simply
doesn't know that 01Dec2000 isn't supposed to be a floating-point number.
> But doesn't a value of 1.0
> result from scanning 01eec2000?
The FP number resulting from this is irrelevant --- if the user actually
asked for this to be read as a time/date string, it'll be overwritten
later by the FP number representing 01Dec2000 in gnuplot's internal time
format. If the user didn't specify time/date format for this column, he
deserves whatever nonsense he may get.
> Now, I see there are some strings to contend with. Aside from that,
> what if one initially goes through the string, using string searching
> routines to look for "dDqQ" and convert them.
No. This would kill both the scanf() format string allowed to be passed
in, if that contained some random 'q' characters, and, perhaps more
importantly, Ethan's datastrings stuff.
|
|
From:
<br...@ph...> - 2005-06-09 13:16:48
|
Robert Hart wrote:
> strtod() is indeed significantly faster than sscanf.
And I still don't see why that should be the case...
> Can anybody come up with any cases where these two methods give
> different results/error handling?
Well, here's an ancient comment right taken directly from datafile.c:
/* cannot trust strtod - eg strtod("-",&p) */
> strtod is ANSI C, so should be
> portable.
So should sscanf().
|
|
From: Robert H. <en...@no...> - 2005-06-09 11:14:38
|
On Thu, 2005-06-09 at 15:26 +0800, Hans-Bernhard Br=F6ker wrote: > That's exactly the error checking I'm talking about. Without it,=20 > handling missing or malformed data would be impossible. > > There's not a lot you could do with strtod() that sscanf() couldn't do=20 > just as well. strtod() is indeed significantly faster than sscanf.=20 I am testing using the following: #ifdef USE_STRTOD char *fin; df_column[df_no_cols].datum=3Dstrtod(s,&fin); used=3Ds-fin; count=3Dused?1:0; #else count =3D sscanf(s, "%lf%n", &df_column[df_no_cols].datum, &used); #endif Can anybody come up with any cases where these two methods give different results/error handling? strtod is ANSI C, so should be portable. I've attached a "non-intrusive" patch, however if this approach is acceptable, I think the OSK path should be removed, and the NO_FORTRAN_NUMS code should be removed or simplified. Are there any cases when somebody wouldn't want a fortran float interpreting properly? Rob =20 --=20 Robert Hart <en...@no...> University of Nottingham This message has been checked for viruses but the contents of an attachment may still contain software viruses, which could damage your computer system: you are advised to perform your own checks. Email communications with the University of Nottingham may be monitored as permitted by UK legislation. |
|
From: Daniel J S. <dan...@ie...> - 2005-06-09 09:13:51
|
I've looked at the current code a bit. It's quite involved, isn't it? F=
or=20
Dimitris' sake, there seems to be some details he may have not accounted =
for and=20
would be difficult to add to the patch. So, would it pay for him to cont=
inue=20
with the patch he's supplied? Or does it look like the current code will=
have=20
to be the starting point?
Let me point out one thing here. With the introduction of binary support=
, I=20
seem to have rewritten df_read_matrix() so that it stores the *whole* set=
of=20
data as floats. In that way, some other features could be reused. Now, =
I don't=20
believe this is a problem as far as time. Neither should it be a memory =
problem=20
(something that would really slow down Dimitris' machine perahps). The b=
lock of=20
memory it consumes will not be much bigger than is eventually stored in m=
emory.
But, what could be a problem--in light of the discussion about FORTRAN=20
doubles--is that instead of "floats" maybe this routine should be storing=
=20
doubles. Certainly a plotting program doesn't need the resolution of dou=
bles.=20
However, perhaps the data that is entered simply needs to have a very, ve=
ry=20
large magnitude exponent and that is why people requested it. (Of course=
, there=20
are ways around that, but...)
So think that over folks. Should it be
static double *
df_read_matrix(int *rows, int *cols)
{
<snip>
double *linearized_matrix =3D NULL;
etc.? Because, if someone uses doubles and, infact, has numbers outside =
the=20
dynamic range of a "float", it's a problem. Do doubles make sense furthe=
r down=20
the line after df_read_matrix?
(more below)
Hans-Bernhard Br=F6ker wrote:
>> There is no error checking in the current code either, unless you
>> mean the count returned by sscanf.=20
>=20
>=20
> That's exactly the error checking I'm talking about. Without it,=20
> handling missing or malformed data would be impossible.
Unless one writes a more sophisticated parser. Given how much testing is=
=20
already done here, that would almost be as good an option. (But perhaps =
not=20
preferable.)
>=20
> > For true error checking we would
>=20
>> need strtod, unless I've overlooked some entirely different method.
>=20
>=20
> There's not a lot you could do with strtod() that sscanf() couldn't do=20
> just as well.
>=20
>>> sscanf() and its %n format are used for a reason.
>=20
>=20
>> But %n is non-portable, as documented by the OSK comment=20
>=20
>=20
> Just because some silly implementation is buggy doesn't exactly mean %n=
=20
> is unportable. %n is as portable as you can ever hope to be: it's an=20
> ANSI requirement.
>=20
>> I'm also curious about that convoluted series of tests added by
>> Corey Satten and labelled "optimization". I'll ask him if he
>> remembers what sort of problem case or test suite he was using.
>=20
>=20
> I dimly remember this being about not running sscanf() on all columns o=
f=20
> a multi-column data file if no extended using specs are in use. If you
> have "using 1:2:3", you can simply ignore columns 4 to 1000. If you=20
> have "using 1:2:(column(some_function($3)), datafile.c has to convert=20
> all 1000 of them.
I don't know how deep into those tests is the normal execution every time=
, but=20
too far and it's inefficient.
There are three tests here that can be done *before* reading data and com=
bined=20
into a single test within tokenise:
(fast_columns =3D=3D 0)
(df_no_use_specs =3D=3D 0)
df_no_use_specs > 5
and somehow my intuition tells me the following could be done in a better=
way:
|| ((df_no_use_specs > 0)
&& (use_spec[0].column =3D=3D dfncp1
|| (df_no_use_specs > 1
&& (use_spec[1].column =3D=3D dfncp1
|| (df_no_use_specs > 2
&& (use_spec[2].column =3D=3D dfncp1
|| (df_no_use_specs > 3
&& (use_spec[3].column =3D=3D dfncp1
|| (df_no_use_specs > 4 && (use_spec[4].column =3D=3D dfncp1 || )
Going from what Hans said, is there some logic that could be done beforeh=
and to=20
set up a more efficient chunk of code?
NEXT ISSUE:
There is this chunk of code:
if (count =3D=3D 1 &&
(s[used] =3D=3D 'd' || s[used] =3D=3D 'D' ||
s[used] =3D=3D 'q' || s[used] =3D=3D 'Q')) {
/* HBB 20001221: avoid breaking parsing of time/date
* strings like 01Dec2000 that would be caused by
* overwriting the 'D' with an 'e'... */
char save_char =3D s[used];
/* might be fortran double */
s[used] =3D 'e';
/* and try again */
count =3D sscanf(s, "%lf", &df_column[df_no_cols].datum);
s[used] =3D save_char;
}
Isn't this a bug? If so, why not? Certainly the date string 01Dec2000 s=
hould=20
not be interpretted as a double. But doesn't a value of 1.0 result from=20
scanning 01eec2000?
I would think that one would have to test for a "+-0123456789" before and=
after=20
a "dDqQ". Then one could feel free to change that character to an 'e' an=
d not=20
worry about it.
Now, I see there are some strings to contend with. Aside from that, what=
if one=20
initially goes through the string, using string searching routines to loo=
k for=20
"dDqQ" and convert them. (Need a simple method to make sure they are not=
within=20
quotes, so one would probably have to include the double quote character =
in the=20
search somehow.) Would it then be possible to use a much simpler routine=
? Or=20
are we still left with using sscanf(), and that is the slow thing here?
Ethan, if
set datafile {no}fortran_floats
is implemented, perhaps a simple test for #{dDqQ}# or #{eE}# on the first=
number=20
in the file would be useful for testing conflicts and giving a warning me=
ssage.=20
(I'm assuming it isn't valid to mix FORTRAN and C format in one file.)
Dan
|
|
From:
<br...@ph...> - 2005-06-09 07:25:02
|
Ethan Merritt wrote: > On Wednesday 08 June 2005 09:47 pm, HBB wrote: >>More to the point, atof has *no* error handling capabilities, >>which I don't believe to be a good idea in this case. > There is no error checking in the current code either, unless you > mean the count returned by sscanf. That's exactly the error checking I'm talking about. Without it, handling missing or malformed data would be impossible. > For true error checking we would > need strtod, unless I've overlooked some entirely different method. There's not a lot you could do with strtod() that sscanf() couldn't do just as well. >>sscanf() and its %n format are used for a reason. > But %n is non-portable, as documented by the OSK comment Just because some silly implementation is buggy doesn't exactly mean %n is unportable. %n is as portable as you can ever hope to be: it's an ANSI requirement. > I'm also curious about that convoluted series of tests added by > Corey Satten and labelled "optimization". I'll ask him if he > remembers what sort of problem case or test suite he was using. I dimly remember this being about not running sscanf() on all columns of a multi-column data file if no extended using specs are in use. If you have "using 1:2:3", you can simply ignore columns 4 to 1000. If you have "using 1:2:(column(some_function($3)), datafile.c has to convert all 1000 of them. |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-09 05:34:06
|
On Wednesday 08 June 2005 09:47 pm, HBB wrote: I assume Robert benchmarked atof because that's what is currently in the source code as a compile-time alternative. But for a real change-over in the default code I think we would need to go with strtod instead. > Ethan Merritt wrote: > > > > I think I'll remove that OSK chunk altogether, > > Please don't. It's quite harmless as it is --- it'll only be used on > one platform, and you would risk breaking that platform completely by > doing this. I don't think it's likely to break anything. The problem is in the %n format, which we can do without by using strtod instead. But yeah, it's only a few lines of code. > More to the point, atof has *no* error handling capabilities, > which I don't believe to be a good idea in this case. There is no error checking in the current code either, unless you mean the count returned by sscanf. For true error checking we would need strtod, unless I've overlooked some entirely different method. > sscanf() and its %n format are used for a reason. But %n is non-portable, as documented by the OSK comment and also by the linux man pages. So if Robert can benchmark the performance of strtod as no worse than the current default code, then replacing both paths with a strtod call may solve all these issues at one stroke. I'm also curious about that convoluted series of tests added by Corey Satten and labelled "optimization". I'll ask him if he remembers what sort of problem case or test suite he was using. We have many more plot options now than when this optimization was done, and it could well be either that it has become pointless or that it is still benificial but has become incomplete. Either way I'd like to know what the original rationale was. -- Ethan A Merritt Biomolecular Structure Center University of Washington 98195-7742 |
|
From:
<br...@ph...> - 2005-06-09 04:46:23
|
Ethan Merritt wrote:
> On Wednesday 08 June 2005 03:53 pm, Robert Hart wrote:
> I think I'll remove that OSK chunk altogether,
Please don't. It's quite harmless as it is --- it'll only be used on
one platform, and you would risk breaking that platform completely by
doing this.
> and make the NO_FORTRAN_NUMS a run-time option. Probably
> set datafile {no}fortran_floats
OK.
>>In the NO_FORTRANS_NUMS code path, atof is used to get the next
>>float. This is much faster than sscanf.
Is it? Why would that be? More to the point, atof has *no* error
handling capabilities, which I don't believe to be a good idea in this case.
> Are you saying that the speed gain could be achieved just by
> replacing sscanf with atof in the standard code path,
That wouldn't work at all --- atof leaves you with no way of knowing the
end of that number, no way to position to the next one, and no way to
find the 'D' in floating point number. sscanf() and its %n format are
used for a reason.
|
|
From:
<br...@ph...> - 2005-06-09 04:04:57
|
Robert Hart wrote: > Questions: Where did this "NO_FORTRAN_NUMS" option come from, From discussions a while back, which I now only dimly remember. The Fortran numbers parsing was added because people insisted they needed it. But it has a potential conflict with certain date/time formats (those that have "Dec" immediately following a number, which was being overwritten to "eec" by the attempt to parse it as a Fortran number), and it makes things slow. > and why isn't it enabled? Because that would break reading of datafiles that people care about. > Is Fortran number support useful? Yes. > Could it be provided in an alternative way (perhaps a run time > option)? Probably. |
|
From: KITA T. <t-...@cc...> - 2005-06-08 23:26:29
|
> From: Hans-Bernhard Br=F6ker <br...@ph...> > Subject: Re: label offset error on 4.1 > Date: Tue, 07 Jun 2005 16:55:42 +0200 > Message-ID: <42A...@ph...> > = > > KITA Toshihiro wrote: > > = > > > The problem happens when I set a label string and the offset at t= he same time. > > > = > > > gnuplot> set xla "x" > > > gnuplot> set xla -10, -2 > > > = > > > is OK, but, I get > > > = > > > gnuplot> set xla "x" -10, -2 > > > ^ > > > invalid expression = > > = > > Please compare this command to the actual syntax description found = under = > > "help xlabel": you're missing the 'offset' keyword. > = > Oops, that was literally 'offset'! > Thanks a lot for your quick reply to my misunderstanding. Just a trivial comment : If the word 'offset' is always necessary before the offset values, I guess the help message for 'help set xlabel' = -----------------------------------------------------------------------= --- ... `character` coordinate system is used. For example, "`set xlabel -1,0= `" will change only the x offset of the title, moving the label roughly o= ne character width to the left. The size of a character depends on both = the ... -----------------------------------------------------------------------= --- should be = -----------------------------------------------------------------------= --- ... `character` coordinate system is used. For example, "`set xlabel offs= et -1,0`" ... -----------------------------------------------------------------------= --- -- = KITA Toshihiro http://t-kita.net/ |
|
From: Robert H. <en...@no...> - 2005-06-08 23:24:23
|
On Wed, 8 Jun 2005, Ethan Merritt wrote: > Are you saying that the speed gain could be achieved just by > replacing sscanf with atof in the standard code path, even though it still > takes the time to check for d/D/q/Q and rescan if found? Not quite. atof is definitely faster than sscanf, however, atof doesn't tell you how many characters it actually read, so that is why (I presume) that sscanf is used instead. However, in both cases, we manually scan for the next seperator one character at a time, so it wouldn't be hard to 'notice' is we pass a dDqQ on the way. Another option I haven't tried/benchmarked is to use strtof. If this is as fast as atof, then we should use it, because it'll give us a pointer to where it finished. Anyway, bed time here, but I think any code that can be culled from this function has got to be a win. Good luck Rob This message has been checked for viruses but the contents of an attachment may still contain software viruses, which could damage your computer system: you are advised to perform your own checks. Email communications with the University of Nottingham may be monitored as permitted by UK legislation. |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-08 23:03:22
|
On Wednesday 08 June 2005 03:53 pm, Robert Hart wrote:
> > The question is whether this comment, which dates back at least to 1999,
> > is in fact correct. Could you please compare the previous benchmarks
> > to the case where the code is prefixed by: #define OSK 1
>
> This is much much worse. (Taking nearly 5 minutes on my benchmark)
OK. No surprise, but I thought it worth testing.
I think I'll remove that OSK chunk altogether, and make the NO_FORTRAN_NUMS
a run-time option. Probably
set datafile {no}fortran_floats
> In the standard code path, sscanf is used to get the next float out of the
> input. Then, if the *NEXT CHARACTER* is a d, D, q, or Q, that character is
> replaced with an "e" and the sscanf is repeated.
>
> In the NO_FORTRANS_NUMS code path, atof is used to get the next
> float. This is much faster than sscanf.
Are you saying that the speed gain could be achieved just by
replacing sscanf with atof in the standard code path, even though it still
takes the time to check for d/D/q/Q and rescan if found?
In that case, maybe we don't even need the run-time option.
thanks,
Ethan
--
Ethan A Merritt merritt@u.washington.edu
Biomolecular Structure Center
Mailstop 357742
University of Washington, Seattle, WA 98195
|
|
From: Robert H. <en...@no...> - 2005-06-08 22:55:18
|
On Wed, 8 Jun 2005, Ethan Merritt wrote: > If you are willing, could you run one more check? > This entire section of code, with or without NO_FORTRAN_NUMS, is > inside a larger block which starts with the comment: > > #ifdef OSK > /* apparently %n does not work. This implementation > * is just as good as the non-OSK one, but close > * to a release (at last) we make it os-9 specific > */ > int count; > char *p = strpbrk(s, "dqDQ"); > if (p != NULL) > *p = 'e'; > > count = sscanf(s, "%lf", &df_column[df_no_cols].datum); > #else > [Previously analysed code is here in the #else] > > The question is whether this comment, which dates back at least to 1999, > is in fact correct. Could you please compare the previous benchmarks > to the case where the code is prefixed by: #define OSK 1 This is much much worse. (Taking nearly 5 minutes on my benchmark) Here's why: In the standard code path, sscanf is used to get the next float out of the input. Then, if the *NEXT CHARACTER* is a d, D, q, or Q, that character is replaced with an "e" and the sscanf is repeated. In the NO_FORTRANS_NUMS code path, atof is used to get the next float. This is much faster than sscanf. In the OSK code path, strpbrk is used to scan *THE ENTIRE INPUT LINE* for the first occurence of d, D, q, or Q. This is *REPEATED* for every value on the line. sscanf is used to read the values which is slow. Rob This message has been checked for viruses but the contents of an attachment may still contain software viruses, which could damage your computer system: you are advised to perform your own checks. Email communications with the University of Nottingham may be monitored as permitted by UK legislation. |
|
From: Ethan M. <merritt@u.washington.edu> - 2005-06-08 22:01:32
|
On Wednesday 08 June 2005 01:50 pm, Robert Hart wrote:
> I used callgrind (and kcachegrind) to profile gnuplot whilst loading
> this datafile. Turned out >95% of time was sping in the sscanf on
> datafile.c:759
Thank you for taking the time to pin this down.
> Questions: Where did this "NO_FORTRAN_NUMS" option come from, and why
> isn't it enabled? Is Fortran number support useful? Could it be provided
> in an alternative way (perhaps a run time option)?
Excellent questions.
Here's a brief excerpt from a Fortran manual:
A double precision constant has the same form as a scaled real constant
except that the E is replaced by D. Examples:
6.1D2 is equivalent to 610.0
+2.3D3 is equivalent to 2300.0
-3.5D-1 is equivalent to -0.35
+4D4 is equivalent to 40000
I have no idea how common this might be in real life data files
fed to gnuplot.
If you are willing, could you run one more check?
This entire section of code, with or without NO_FORTRAN_NUMS, is
inside a larger block which starts with the comment:
#ifdef OSK
/* apparently %n does not work. This implementation
* is just as good as the non-OSK one, but close
* to a release (at last) we make it os-9 specific
*/
int count;
char *p = strpbrk(s, "dqDQ");
if (p != NULL)
*p = 'e';
count = sscanf(s, "%lf", &df_column[df_no_cols].datum);
#else
[Previously analysed code is here in the #else]
The question is whether this comment, which dates back at least to 1999,
is in fact correct. Could you please compare the previous benchmarks
to the case where the code is prefixed by: #define OSK 1
I know this will break the checks for separators other than whitespace,
but we can sort that out afterwards. It would be nice to get a speed-up
of 10X by deleting 80+ lines of obfuscated code :-)
--
Ethan A Merritt merritt@u.washington.edu
Biomolecular Structure Center
Mailstop 357742
University of Washington, Seattle, WA 98195
|
|
From: Robert H. <en...@no...> - 2005-06-08 20:50:57
|
On Wed, 2005-06-08 at 15:31 +0300, Dimitrios Apostolou wrote: > Did anyone actually tried it to see the speed improvement? I have done > no benchmarks but what I described in an earlier email (how to plot big > SRTM files) now works *much* faster (*10 or more speed improvement). I've had a look into this using n x n ascii matrix (generated from perl's sin() function) Existing code: 66 seconds after adding: #define NO_FORTRAN_NUMS to top of datafile.c: 5.9 seconds. Rationale: I used callgrind (and kcachegrind) to profile gnuplot whilst loading this datafile. Turned out >95% of time was sping in the sscanf on datafile.c:759 Questions: Where did this "NO_FORTRAN_NUMS" option come from, and why isn't it enabled? Is Fortran number support useful? Could it be provided in an alternative way (perhaps a run time option)? Rob |
|
From: Daniel J S. <dan...@ie...> - 2005-06-08 18:52:48
|
Dimitrios Apostolou wrote: > For now I will be happy to see an entry in the TODO file about > optimizing the highly innefficient "matrix" parser. In the future, if > you haven't found the time to fix it, perhaps I will submit a proper > patch, compliant to the coding standards and good enough for you to use it. You are familiar with the SourceForge "patch" site, aren't you? What you are proposing is a slightly bigger fix and hence requires some consideration. Often smaller patches submitted through discussion can go to CVS as a no-brainer. However, in the case of larger patches, SourceForge is nice for maintaining gradual development, both for other developers and yourself. You can come back to it whenever you have the free time. (Consider creating a SourceForge account.) Dan |