You can subscribe to this list here.
| 2001 |
Jan
|
Feb
(1) |
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
|
Sep
|
Oct
|
Nov
|
Dec
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2002 |
Jan
(1) |
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
|
Nov
(1) |
Dec
|
| 2003 |
Jan
|
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
(83) |
Nov
(57) |
Dec
(111) |
| 2004 |
Jan
(38) |
Feb
(121) |
Mar
(107) |
Apr
(241) |
May
(102) |
Jun
(190) |
Jul
(239) |
Aug
(158) |
Sep
(184) |
Oct
(193) |
Nov
(47) |
Dec
(68) |
| 2005 |
Jan
(190) |
Feb
(105) |
Mar
(99) |
Apr
(65) |
May
(92) |
Jun
(250) |
Jul
(197) |
Aug
(128) |
Sep
(101) |
Oct
(183) |
Nov
(186) |
Dec
(42) |
| 2006 |
Jan
(102) |
Feb
(122) |
Mar
(154) |
Apr
(196) |
May
(181) |
Jun
(281) |
Jul
(310) |
Aug
(198) |
Sep
(145) |
Oct
(188) |
Nov
(134) |
Dec
(90) |
| 2007 |
Jan
(134) |
Feb
(181) |
Mar
(157) |
Apr
(57) |
May
(81) |
Jun
(204) |
Jul
(60) |
Aug
(37) |
Sep
(17) |
Oct
(90) |
Nov
(122) |
Dec
(72) |
| 2008 |
Jan
(130) |
Feb
(108) |
Mar
(160) |
Apr
(38) |
May
(83) |
Jun
(42) |
Jul
(75) |
Aug
(16) |
Sep
(71) |
Oct
(57) |
Nov
(59) |
Dec
(152) |
| 2009 |
Jan
(73) |
Feb
(213) |
Mar
(67) |
Apr
(40) |
May
(46) |
Jun
(82) |
Jul
(73) |
Aug
(57) |
Sep
(108) |
Oct
(36) |
Nov
(153) |
Dec
(77) |
| 2010 |
Jan
(42) |
Feb
(171) |
Mar
(150) |
Apr
(6) |
May
(22) |
Jun
(34) |
Jul
(31) |
Aug
(38) |
Sep
(32) |
Oct
(59) |
Nov
(13) |
Dec
(62) |
| 2011 |
Jan
(114) |
Feb
(139) |
Mar
(126) |
Apr
(51) |
May
(53) |
Jun
(29) |
Jul
(41) |
Aug
(29) |
Sep
(35) |
Oct
(87) |
Nov
(42) |
Dec
(20) |
| 2012 |
Jan
(111) |
Feb
(66) |
Mar
(35) |
Apr
(59) |
May
(71) |
Jun
(32) |
Jul
(11) |
Aug
(48) |
Sep
(60) |
Oct
(87) |
Nov
(16) |
Dec
(38) |
| 2013 |
Jan
(5) |
Feb
(19) |
Mar
(41) |
Apr
(47) |
May
(14) |
Jun
(32) |
Jul
(18) |
Aug
(68) |
Sep
(9) |
Oct
(42) |
Nov
(12) |
Dec
(10) |
| 2014 |
Jan
(14) |
Feb
(139) |
Mar
(137) |
Apr
(66) |
May
(72) |
Jun
(142) |
Jul
(70) |
Aug
(31) |
Sep
(39) |
Oct
(98) |
Nov
(133) |
Dec
(44) |
| 2015 |
Jan
(70) |
Feb
(27) |
Mar
(36) |
Apr
(11) |
May
(15) |
Jun
(70) |
Jul
(30) |
Aug
(63) |
Sep
(18) |
Oct
(15) |
Nov
(42) |
Dec
(29) |
| 2016 |
Jan
(37) |
Feb
(48) |
Mar
(59) |
Apr
(28) |
May
(30) |
Jun
(43) |
Jul
(47) |
Aug
(14) |
Sep
(21) |
Oct
(26) |
Nov
(10) |
Dec
(2) |
| 2017 |
Jan
(26) |
Feb
(27) |
Mar
(44) |
Apr
(11) |
May
(32) |
Jun
(28) |
Jul
(75) |
Aug
(45) |
Sep
(35) |
Oct
(285) |
Nov
(99) |
Dec
(16) |
| 2018 |
Jan
(8) |
Feb
(8) |
Mar
(42) |
Apr
(35) |
May
(23) |
Jun
(12) |
Jul
(16) |
Aug
(11) |
Sep
(8) |
Oct
(16) |
Nov
(5) |
Dec
(8) |
| 2019 |
Jan
(9) |
Feb
(28) |
Mar
(4) |
Apr
(10) |
May
(7) |
Jun
(4) |
Jul
(4) |
Aug
|
Sep
(4) |
Oct
|
Nov
(23) |
Dec
(3) |
| 2020 |
Jan
(19) |
Feb
(3) |
Mar
(22) |
Apr
(17) |
May
(10) |
Jun
(69) |
Jul
(18) |
Aug
(23) |
Sep
(25) |
Oct
(11) |
Nov
(20) |
Dec
(9) |
| 2021 |
Jan
(1) |
Feb
(7) |
Mar
(9) |
Apr
|
May
(1) |
Jun
(8) |
Jul
(6) |
Aug
(8) |
Sep
(7) |
Oct
|
Nov
(2) |
Dec
(23) |
| 2022 |
Jan
(23) |
Feb
(9) |
Mar
(9) |
Apr
|
May
(8) |
Jun
(1) |
Jul
(6) |
Aug
(8) |
Sep
(30) |
Oct
(5) |
Nov
(4) |
Dec
(6) |
| 2023 |
Jan
(2) |
Feb
(5) |
Mar
(7) |
Apr
(3) |
May
(8) |
Jun
(45) |
Jul
(8) |
Aug
|
Sep
(2) |
Oct
(14) |
Nov
(7) |
Dec
(2) |
| 2024 |
Jan
(4) |
Feb
(4) |
Mar
|
Apr
(7) |
May
(2) |
Jun
(1) |
Jul
|
Aug
(5) |
Sep
|
Oct
|
Nov
(4) |
Dec
(14) |
| 2025 |
Jan
(22) |
Feb
(6) |
Mar
(5) |
Apr
(14) |
May
(6) |
Jun
(11) |
Jul
(19) |
Aug
|
Sep
(17) |
Oct
(1) |
Nov
(2) |
Dec
(18) |
| 2026 |
Jan
|
Feb
|
Mar
(5) |
Apr
|
May
(2) |
Jun
(1) |
Jul
(6) |
Aug
(1) |
Sep
|
Oct
|
Nov
|
Dec
|
|
From: <pl...@pi...> - 2016-06-16 20:05:16
|
On 16/06/16 19:44, Ethan A Merritt wrote: > On Thursday, 16 June, 2016 18:00:00 pl...@pi... wrote: >> On 16/06/16 17:23, sfeam wrote: >> >> >>> - The comparison is to a string, not a numerical value, so -99.00 ne -99.0 ne -99 >> Sounds reasonable. From the gnuplot POV it is a missing *string* ; if >> the cvs is output by a spreadsheet or other software it seems reasonable >> to expect consistent string formatting ( although Excel could have >> different cell formats, that is probably too much to try and anticipate. >> >> >> >- Leading whitespace is ignore but trailing whitespace is not. >> >> Seems inconsistent. Was this a programming convenience for minimal >> coding changes or is there a functional logic behind this? > > I though the consensus from a couple of days ago was that any difference > in the remainder of the field was significant, hence extra trailing > characters would mean that the match was imperfect. > Previously "missing A" and "missing B" were both matched as "missing". > Now they are not, even if A is a <tab> or '\n' or '\r'. I don't know what the consensus was but my comment on that was that trailing WS should be stripped, as it is with leading WS. I was suggesting that any non-WS following the missing string meant the match failed. Specifically relating to your " ignore A" case. I did not suggest WS"ignore"WS should fail. If leading space is stripped, I'm not sure I see why trailing is not also stripped. > >> > ... and requires that the next character is a field-terminator. >> >> I presume field-terminator.means FS or EOL. > > Separator or null. > EOL is legal with in a csv field, although if you have such a file good > luck to you. When a line of data is read in to gnuplot it is transferred > to a null-terminated string, so the check for null should catch the true > end-of-line. > >> Does this cater for WS at end of line without an explicit FS, >> or does this fall foul of previous point? > > You mean like a DOS-style file with <cr><nl> at the end of the line? > So far as I know this is properly handled by stripping away both line > termination characters on input. But more testing wouldn't hurt. > > Ethan > No , I was not talking about CRLF end of line. It is quite common to have 'invisible' WS after the last field and being the last field probably no FS. This is especially the case if there was a comment : WS to provide visual separation or align comments: 1,2,3,-999 # last column data got lost in paper records ! Once the # is replaced by #0 to truncate out the comment , this line would fall foul of your new scheme I think. I see no real reason not to remove the tailing WS , it seems a little odd to strip one end an not the other. CSV is pretty illegible at the best of times. If need to dump a spreadsheet to CSV I often separate with " , " to make the result a little easier to read afterwards. Unless I'm missing something , I don't see any reason or advantage to not stripping trailing WS. Not wishing to be finicky, but you seemed interesting is considering any corner cases. Peter. > > > > Ethan > |
|
From: Ethan A M. <sf...@us...> - 2016-06-16 18:46:21
|
On Thursday, 16 June, 2016 18:00:00 pl...@pi... wrote: > On 16/06/16 17:23, sfeam wrote: > > > > - The comparison is to a string, not a numerical value, so -99.00 ne -99.0 ne -99 > Sounds reasonable. From the gnuplot POV it is a missing *string* ; if > the cvs is output by a spreadsheet or other software it seems reasonable > to expect consistent string formatting ( although Excel could have > different cell formats, that is probably too much to try and anticipate. > > > >- Leading whitespace is ignore but trailing whitespace is not. > > Seems inconsistent. Was this a programming convenience for minimal > coding changes or is there a functional logic behind this? I though the consensus from a couple of days ago was that any difference in the remainder of the field was significant, hence extra trailing characters would mean that the match was imperfect. Previously "missing A" and "missing B" were both matched as "missing". Now they are not, even if A is a <tab> or '\n' or '\r'. > > ... and requires that the next character is a field-terminator. > > I presume field-terminator.means FS or EOL. Separator or null. EOL is legal with in a csv field, although if you have such a file good luck to you. When a line of data is read in to gnuplot it is transferred to a null-terminated string, so the check for null should catch the true end-of-line. > Does this cater for WS at end of line without an explicit FS, > or does this fall foul of previous point? You mean like a DOS-style file with <cr><nl> at the end of the line? So far as I know this is properly handled by stripping away both line termination characters on input. But more testing wouldn't hurt. Ethan Ethan |
|
From: <pl...@pi...> - 2016-06-16 17:35:26
|
On 16/06/16 17:23, sfeam wrote: > - The comparison is to a string, not a numerical value, so -99.00 ne -99.0 ne -99 Sounds reasonable. From the gnuplot POV it is a missing *string* ; if the cvs is output by a spreadsheet or other software it seems reasonable to expect consistent string formatting ( although Excel could have different cell formats, that is probably too much to try and anticipate. >- Leading whitespace is ignore but trailing whitespace is not. Seems inconsistent. Was this a programming convenience for minimal coding changes or is there a functional logic behind this? > ... and requires that the next character is a field-terminator. I presume field-terminator.means FS or EOL. Does this cater for WS at end of line without an explicit FS, or does this fall foul of previous point? Thanks. Peter. |
|
From: sfeam <sf...@us...> - 2016-06-16 16:24:15
|
On Wednesday, 15 June 2016 11:27:27 AM pl...@pi... wrote: > On 15/06/16 00:06, Ethan A Merritt wrote: > >> I think you are correct that there is a bug in this part of the code. > >> > > The full-length string from 'set missing' is tested against the > >> > > start of the field contents (after removing leading whitespace); > >> > > then the subsequent character is tested to see if it is whitespace > >> > > rather than a continuation of whatever string is in the field. > >> > > So it works with a tab-separated *.csv file because <tab> counts > >> > > as whitespace, but fails with a comma-separated file because the > >> > > comma is mis-interpreted as part of the field content. > >> > > further thoughts: > > The cause of this then, seems to be an oversight. The end of the string > is being tested as though it was the default case of WSpace separators > and not the specified separator. > > It is a little irrelevant what name is give to this sort of file, the > key point is that the check you describe is not using the current > datafile separator. > > Presumably the same thing would happen if someone had a file using colon > ( or any other non WS char ) as separator and had correctly specified it > with > > > set datafile separator ":" > > I have not tested this explicitly but there is nothing special about > using comma sep. so I presume the same bug would manifest. > > " the subsequent character is tested to see if it is whitespace" > > It seems that this test should be firstly a test for 'separator' and > then additionally for white-space + separator. As previously stated > "ignore A" probably should count as a match. Substrings counting as a > match is not described anywhere and I see not reason for this to be > taken as a hit. > > Thanks for looking into this. > Peter. I have made a change to datafile.c:check_missing() in CVS for both 5.0 and 5.1. In the case of a csv file (i.e. "set datafile separator" is non-blank) it now checks for a match of the field contents to the "missing" string and requires that the next character is a field-terminator. Notes: - Leading whitespace is ignore but trailing whitespace is not. - This is a obviously a change, so possibly there are existing scripts that break. - The comparison is to a string, not a numerical value, so -99.00 ne -99.0 ne -99 - If the "missing" string is quoted in the data file it will not be recognized. - If the "missing" string itself contains quotes, the behaviour is not specified This change does not include an earlier suggestion to provide an option that causes NaN (not-a-number) values to be treated as missing data. I am inclined to add this also, probably as a new keyword "set datafile missing NaN". See `help missing` for detail on the current handling of missing and NaN values. Ethan |
|
From: Philipp K. J. <ja...@ie...> - 2016-06-16 15:01:43
|
[snip] > > If you have selected a range manually ( mouse selected zoom ) the > third line will not remove the implicit ranges you have set. If you > then want to plot [:*] "file" u 1:2 , you will have to do so > explicitly to remove the range settings applied with the mouse. > > Could that be the source of the 'inconsistency' ? Ah - interesting hint! I forgot to mention in my original email that I never set plot ranges with the mouse, so that's not it. But your suggestion made me think of something else: occasionally I use the cursor-keys when the plot window has focus (which moves the plot). I only do this by accident (because I don't find that feature useful). But apparently, doing so does an implied "set xrange". This behavior is reproducible, so that's probably the explanation for what I have seen. Thanks for the hint. Best, Ph. |
|
From: <pl...@pi...> - 2016-06-16 10:51:57
|
On 16/06/16 03:34, Philipp K. Janert wrote: > > I have observed an intermittent (not reliably repeatable) > problem with inline plot ranges. Has anybody else seen this? > > Basically, some times plot ranges seem to get "stuck", > rather than defaulting back to [*:*]. > > Consider the following sequence of commands: > > plot "file" u 1:2 > plot [:100] "file" u 1:2 > plot "file" u 1:2 > > I'd expect the third call to produce the same plot as the > first. Most of the time this is true, but sometimes the > third call produces the same result as the second, and I > need to use plot [:*] "file" u 1:2. > > I have also observed this behavior for the vertical axis. > > I have not been able to identify the circumstances that > lead to this behavior. > > Has anybody seen something similar or some hypotheses, > that might help narrow down on the conditions that > produce this behavior? > > Best, > > Ph. > If you have selected a range manually ( mouse selected zoom ) the third line will not remove the implicit ranges you have set. If you then want to plot [:*] "file" u 1:2 , you will have to do so explicitly to remove the range settings applied with the mouse. Could that be the source of the 'inconsistency' ? Peter. |
|
From: <pl...@pi...> - 2016-06-16 08:01:59
|
On 15/06/16 20:11, sfeam wrote: > On Wednesday, 15 June 2016 06:36:53 PM Tait wrote: >> >>>> The conventional indication of missing data in a *.csv file is simply >>>> an empty field. This obviously is not possible in a whitespace-separated >>>> file. Gnuplot's use of "set missing" is outside any standard practice >>>> I know of for csv files, so anything we choose is likely to strike >>>> someone as wrong. >> >> I don't follow the comment in the second sentence. > > I was contrasting csv files to whitespace-separated files. > > If you have only whitespace to separate values, then you can't simply > omit a field because this is indistinguishable from shifting all the > remaining fields over by one. That's why we need a "missing" > placeholder. In a csv file you shouldn't really need a "missing" > placeholder because an empty field is unambiguous. > > Ethan It should be remembered that csv are rarely a data storage object but simply a text dump of something else. I don't think anyone would chose csv as a working file format. It's more a means of transmission. So the question is what is the source data format that the csv is a dump of. Missing value flags like -999 etc are commonly used in software storing data in numerical arrays where every datum must have a finite value assigned. An array value cannot be 'empty'. A csv dump from such software will most likely preserve the missing value flag rather then strip them out. Even though a spreadsheet can have an empty cell, it is often useful to have an affirmative flag that indicates that the cell has been processed and determined to be a missing datum, rather than just having been overlooked or not yet processed. So even though a csv file can represent an empty value, this does not mean missing value marker is not necessary. Peter. > >> Empty fields in a >> TSV file* are indicated by having no data in between the field >> separators. Not only is it possible, but it's quite intuitive, I >> think. I'm adding spaces for clarity, but those spaces wouldn't be >> in the actual file: >> >> header1 \t header2 \t header3 \t header4 >> data1 \t data2 \t data3 \t data4 >> data5 \t \t data7 \t data8 >> ... >> >> Where data6 would be, is an empty field. >> >> (* as an aside, TSV or "tab-separated text" is the term I always see >> used for tab-separated values. I've never heard of someone refer to >> a tab-separated file as "CSV".) >> >>>> For instance, RFC-4180, the closest thing to a csv standard, states that >>>> "any field may be quoted with double quotes". So in the example above, >>>> should we ignore this line? >>>> 5, 5, "ignore", 5 >>>> This one? >>>> 5, 5, " ignore ", 5 >>> >>> Ok, in the absence of any properly defined standard , where software >>> like Excel ( probably the most common source of "CSV" files for a lot of >>> people ) produces comma separated variables without using commas, it is >>> likely to be messy. >>> >>> 5, 5, "ignore", 5 >>> >>> This seems a bit of a contrived case, what software will quote one field >>> in a line but not the others? >> >> Excel does exactly this. Fields are unquoted in general, but (only) >> if they contain delimiter or quoting characters, then they are >> quoted. If they contain quote characters, quotes are double-quoted. >> Delimiter characters are not just "," for CSV, but also newlines. >> Consider three rows of data, each containing two fields: >> >> row 1: ab cd >> row 2: e\nf g,h >> row 3: i"j k<space>m >> >> Excel will produce a CSV that looks like this: >> >> ab,cd >> "e >> f","g,h" >> "i""j",k m >> >> This is obviously a contrived pathological case, but it's >> illustrative of what common software "out there" might do. >> Of course, backslash-escaping is also a common convention, >> and for the same input, it might produce a CSV like: >> >> ab,cd >> e\ >> f,g\,h >> i"j,k m >> >> As Ethan mentioned, any convention will break some >> expectations/compatibility, unless the plan is to build in >> a wide range of application- or convention-specific input >> filters. (And those filters implemented in Perl is usually >> how I get by and produce the format gnuplot expects.) >> > > |
|
From: Philipp K. J. <ja...@ie...> - 2016-06-16 02:52:37
|
I have observed an intermittent (not reliably repeatable) problem with inline plot ranges. Has anybody else seen this? Basically, some times plot ranges seem to get "stuck", rather than defaulting back to [*:*]. Consider the following sequence of commands: plot "file" u 1:2 plot [:100] "file" u 1:2 plot "file" u 1:2 I'd expect the third call to produce the same plot as the first. Most of the time this is true, but sometimes the third call produces the same result as the second, and I need to use plot [:*] "file" u 1:2. I have also observed this behavior for the vertical axis. I have not been able to identify the circumstances that lead to this behavior. Has anybody seen something similar or some hypotheses, that might help narrow down on the conditions that produce this behavior? Best, Ph. |
|
From: sfeam <sf...@us...> - 2016-06-15 19:12:14
|
On Wednesday, 15 June 2016 06:36:53 PM Tait wrote: > > > > The conventional indication of missing data in a *.csv file is simply > > > an empty field. This obviously is not possible in a whitespace-separated > > > file. Gnuplot's use of "set missing" is outside any standard practice > > > I know of for csv files, so anything we choose is likely to strike > > > someone as wrong. > > I don't follow the comment in the second sentence. I was contrasting csv files to whitespace-separated files. If you have only whitespace to separate values, then you can't simply omit a field because this is indistinguishable from shifting all the remaining fields over by one. That's why we need a "missing" placeholder. In a csv file you shouldn't really need a "missing" placeholder because an empty field is unambiguous. Ethan > Empty fields in a > TSV file* are indicated by having no data in between the field > separators. Not only is it possible, but it's quite intuitive, I > think. I'm adding spaces for clarity, but those spaces wouldn't be > in the actual file: > > header1 \t header2 \t header3 \t header4 > data1 \t data2 \t data3 \t data4 > data5 \t \t data7 \t data8 > ... > > Where data6 would be, is an empty field. > > (* as an aside, TSV or "tab-separated text" is the term I always see > used for tab-separated values. I've never heard of someone refer to > a tab-separated file as "CSV".) > > > > For instance, RFC-4180, the closest thing to a csv standard, states that > > > "any field may be quoted with double quotes". So in the example above, > > > should we ignore this line? > > > 5, 5, "ignore", 5 > > > This one? > > > 5, 5, " ignore ", 5 > > > > Ok, in the absence of any properly defined standard , where software > > like Excel ( probably the most common source of "CSV" files for a lot of > > people ) produces comma separated variables without using commas, it is > > likely to be messy. > > > > 5, 5, "ignore", 5 > > > > This seems a bit of a contrived case, what software will quote one field > > in a line but not the others? > > Excel does exactly this. Fields are unquoted in general, but (only) > if they contain delimiter or quoting characters, then they are > quoted. If they contain quote characters, quotes are double-quoted. > Delimiter characters are not just "," for CSV, but also newlines. > Consider three rows of data, each containing two fields: > > row 1: ab cd > row 2: e\nf g,h > row 3: i"j k<space>m > > Excel will produce a CSV that looks like this: > > ab,cd > "e > f","g,h" > "i""j",k m > > This is obviously a contrived pathological case, but it's > illustrative of what common software "out there" might do. > Of course, backslash-escaping is also a common convention, > and for the same input, it might produce a CSV like: > > ab,cd > e\ > f,g\,h > i"j,k m > > As Ethan mentioned, any convention will break some > expectations/compatibility, unless the plan is to build in > a wide range of application- or convention-specific input > filters. (And those filters implemented in Perl is usually > how I get by and produce the format gnuplot expects.) > |
|
From: Tait <gnu...@t4...> - 2016-06-15 18:35:37
|
> > The conventional indication of missing data in a *.csv file is simply > > an empty field. This obviously is not possible in a whitespace-separated > > file. Gnuplot's use of "set missing" is outside any standard practice > > I know of for csv files, so anything we choose is likely to strike > > someone as wrong. I don't follow the comment in the second sentence. Empty fields in a TSV file* are indicated by having no data in between the field separators. Not only is it possible, but it's quite intuitive, I think. I'm adding spaces for clarity, but those spaces wouldn't be in the actual file: header1 \t header2 \t header3 \t header4 data1 \t data2 \t data3 \t data4 data5 \t \t data7 \t data8 ... Where data6 would be, is an empty field. (* as an aside, TSV or "tab-separated text" is the term I always see used for tab-separated values. I've never heard of someone refer to a tab-separated file as "CSV".) > > For instance, RFC-4180, the closest thing to a csv standard, states that > > "any field may be quoted with double quotes". So in the example above, > > should we ignore this line? > > 5, 5, "ignore", 5 > > This one? > > 5, 5, " ignore ", 5 > > Ok, in the absence of any properly defined standard , where software > like Excel ( probably the most common source of "CSV" files for a lot of > people ) produces comma separated variables without using commas, it is > likely to be messy. > > 5, 5, "ignore", 5 > > This seems a bit of a contrived case, what software will quote one field > in a line but not the others? Excel does exactly this. Fields are unquoted in general, but (only) if they contain delimiter or quoting characters, then they are quoted. If they contain quote characters, quotes are double-quoted. Delimiter characters are not just "," for CSV, but also newlines. Consider three rows of data, each containing two fields: row 1: ab cd row 2: e\nf g,h row 3: i"j k<space>m Excel will produce a CSV that looks like this: ab,cd "e f","g,h" "i""j",k m This is obviously a contrived pathological case, but it's illustrative of what common software "out there" might do. Of course, backslash-escaping is also a common convention, and for the same input, it might produce a CSV like: ab,cd e\ f,g\,h i"j,k m As Ethan mentioned, any convention will break some expectations/compatibility, unless the plan is to build in a wide range of application- or convention-specific input filters. (And those filters implemented in Perl is usually how I get by and produce the format gnuplot expects.) |
|
From: sfeam <sf...@us...> - 2016-06-15 15:32:22
|
On Wednesday, 15 June 2016 08:40:46 AM pl...@pi... wrote: > On 15/06/16 00:06, Ethan A Merritt wrote: > >> Only something which IS the 'missing' string or the string with leading > >> and/or trailing white-space should match, IMO. > > > > The conventional indication of missing data in a *.csv file is simply > > an empty field. This obviously is not possible in a whitespace-separated > > file. Gnuplot's use of "set missing" is outside any standard practice > > I know of for csv files, so anything we choose is likely to strike > > someone as wrong. > > > > For instance, RFC-4180, the closest thing to a csv standard, states that > > "any field may be quoted with double quotes". So in the example above, > > should we ignore this line? > > 5, 5, "ignore", 5 > > This one? > > 5, 5, " ignore ", 5 > > > > > > Ethan > > > >> > >> Peter. > > > > > > > Ok, in the absence of any properly defined standard , where software > like Excel ( probably the most common source of "CSV" files for a lot of > people ) produces comma separated variables without using commas, it is > likely to be messy. > > 5, 5, "ignore", 5 > > This seems a bit of a contrived case, what software will quote one field > in a line but not the others? Excel for one. It depends on what "format type" you assign to the column. > How would gnuplot cope with : > > "5","5", "ignore", "5" Gnuplot explicitly checks for both numerical and quoted numerical input in csv files exactly because of this issue. But the concept of checking for both quoted and unquoted "missing" strings never occurred to me until just now. > Looking at the bug I reported may be a chance to review this whole messy > subject but it seems like a diversion from the clear bug case. > > If gnuplot scans for the position of the field separators, it should be > stopping BEFORE it gets to the next one when testing for occurrences of > the 'missing' string. > > That seems to be a simple bug that does not open a whole can of csv worms. Yeah, but while revisiting the code it seems like a good idea to not only fix the specific case in the bug report but also any other corner cases we can think of. Ethan |
|
From: <pl...@pi...> - 2016-06-15 10:27:38
|
On 15/06/16 00:06, Ethan A Merritt wrote: >> I think you are correct that there is a bug in this part of the code. >> > > The full-length string from 'set missing' is tested against the >> > > start of the field contents (after removing leading whitespace); >> > > then the subsequent character is tested to see if it is whitespace >> > > rather than a continuation of whatever string is in the field. >> > > So it works with a tab-separated *.csv file because <tab> counts >> > > as whitespace, but fails with a comma-separated file because the >> > > comma is mis-interpreted as part of the field content. >> further thoughts: The cause of this then, seems to be an oversight. The end of the string is being tested as though it was the default case of WSpace separators and not the specified separator. It is a little irrelevant what name is give to this sort of file, the key point is that the check you describe is not using the current datafile separator. Presumably the same thing would happen if someone had a file using colon ( or any other non WS char ) as separator and had correctly specified it with set datafile separator ":" I have not tested this explicitly but there is nothing special about using comma sep. so I presume the same bug would manifest. " the subsequent character is tested to see if it is whitespace" It seems that this test should be firstly a test for 'separator' and then additionally for white-space + separator. As previously stated "ignore A" probably should count as a match. Substrings counting as a match is not described anywhere and I see not reason for this to be taken as a hit. Thanks for looking into this. Peter. |
|
From: <pl...@pi...> - 2016-06-15 08:20:29
|
On 15/06/16 00:06, Ethan A Merritt wrote: > On Tuesday, 14 June, 2016 23:00:19 pl...@pi... wrote: >> On 14/06/16 21:00, Ethan A Merritt wrote: >>> On Tuesday, 14 June, 2016 13:10:13 pl...@pi... wrote: >>>> On 14/06/16 12:52, Allin Cottrell wrote: >>>>> On Tue, 14 Jun 2016, pl...@pi... wrote: >>>>> >>>>>> On 14/06/16 09:43, pl...@pi... wrote: >>>>>>> >>>>>>> I have a data csv datafile which uses -99.99 as it missing data value. >>>>>>> >>>>>>> with the following settings the 'missing' data are getting plotted. >>>>>>> >>>>>>> set datafile separator "," >>>>>>> set datafile missing "-99.99" >>>>>>> >>>>>>> >>>>>>> show datafile missing >>>>>>> >>>>>>> "-99.99" in datafile is interpreted as missing value >>>>>>> >>> >>> I think you are correct that there is a bug in this part of the code. >>> The full-length string from 'set missing' is tested against the >>> start of the field contents (after removing leading whitespace); >>> then the subsequent character is tested to see if it is whitespace >>> rather than a continuation of whatever string is in the field. >>> So it works with a tab-separated *.csv file because <tab> counts >>> as whitespace, but fails with a comma-separated file because the >>> comma is mis-interpreted as part of the field content. >> >> Thanks Ethan, >> >> First comment: a tab separated file is not a CSV file. The C mean comma >> separated. > > In practice this is not true. Pretty much any program I know of that > supports *.csv files allows you to specify what character is used as > a field separator. > > Quoting Wikipedia: > > "the term "CSV" also denotes some closely related delimiter-separated > formats that use different field delimiters. These include tab-separated > values and space-separated values. A delimiter that is not present in > the field data (such as tab) keeps the format parsing simple. > These alternate delimiter-separated files are often even given a > .csv extension, despite the use of a non-comma field separator." > >> I would suggest that the correct, structured way to do this is to break >> into fields using the current field separator, then test whether any >> fields match the missing string ( with the white-space caveats ). >> >> If I follow your explanation, it would seem that currently the whole >> line is scanned for the 'missing' string before it is split into fields, >> or it is being parsed twice. > > Not quite. The input line is scanned for field separators, the start > of each field is noted, then it goes back to process them one-by-one. > >> >> 2, 2, ignore A, 2 >> 3, 3, ignore B, 3 >> 4, 4, ignore, 4 >> >> >> IMO 2 and 3 should not match since the field is not equal to the >> 'missing' string but simply contains it. This sounds like asking for >> trouble. Allowing white space seems sensible flexibility on insisting on >> an exact match since it is often added for human readability, as is the >> case here. >> >> Only something which IS the 'missing' string or the string with leading >> and/or trailing white-space should match, IMO. > > The conventional indication of missing data in a *.csv file is simply > an empty field. This obviously is not possible in a whitespace-separated > file. Gnuplot's use of "set missing" is outside any standard practice > I know of for csv files, so anything we choose is likely to strike > someone as wrong. > > For instance, RFC-4180, the closest thing to a csv standard, states that > "any field may be quoted with double quotes". So in the example above, > should we ignore this line? > 5, 5, "ignore", 5 > This one? > 5, 5, " ignore ", 5 > > > Ethan > >> >> Peter. > > Ok, in the absence of any properly defined standard , where software like Excel ( probably the most common source of "CSV" files for a lot of people ) produces comma separated variables without using commas, it is likely to be messy. 5, 5, "ignore", 5 This seems a bit of a contrived case, what software will quote one field in a line but not the others? How would gnuplot cope with : "5","5", "ignore", "5" Looking at the bug I reported may be a chance to review this whole messy subject but it seems like a diversion from the clear bug case. If gnuplot scans for the position of the field separators, it should be stopping BEFORE it gets to the next one when testing for occurrences of the 'missing' string. That seems to be a simple bug that does not open a whole can of csv worms. Peter. |
|
From: Tait <gnu...@t4...> - 2016-06-15 01:05:45
|
*shrug* I don't think I'm qualified to offer any more input than I already have. Doubly so because I'm a user and not a developer of gnuplot, and bear no liability in any case. I am content with whatever decision is made having awareness of the situation. I just wouldn't want anyone to fall astray out of ignorance of the issue. Tait Tatsuro MATSUOKA <tma...@ya...> said (on 2016/06/14): > The change > > 2016-06-14 Bastian Maerkisch <bma...@we...> > > * src/readline.c src/plot.c config/config.mgw config/mingw/Makefile: > Allow console mode gnuplot on Windows to use GNU readline, which is > more powerful than the builtin code, but still does not handle Unicode > input on Windows. > > For GNU readline, we had discussed the below: > > http://gnuplot.10905.n7.nabble.com/readline-to-gnuplot-exe-console-mode-for-windows-td11173.html > > I at the moment do not link the GNU readline for my CVS binary distribution unless > a new statement will appear about the GNU readline. > > Tatsuro |
|
From: Ethan A M. <sf...@us...> - 2016-06-14 23:06:49
|
On Tuesday, 14 June, 2016 23:00:19 pl...@pi... wrote: > On 14/06/16 21:00, Ethan A Merritt wrote: > > On Tuesday, 14 June, 2016 13:10:13 pl...@pi... wrote: > >> On 14/06/16 12:52, Allin Cottrell wrote: > >>> On Tue, 14 Jun 2016, pl...@pi... wrote: > >>> > >>>> On 14/06/16 09:43, pl...@pi... wrote: > >>>>> > >>>>> I have a data csv datafile which uses -99.99 as it missing data value. > >>>>> > >>>>> with the following settings the 'missing' data are getting plotted. > >>>>> > >>>>> set datafile separator "," > >>>>> set datafile missing "-99.99" > >>>>> > >>>>> > >>>>> show datafile missing > >>>>> > >>>>> "-99.99" in datafile is interpreted as missing value > >>>>> > > > > I think you are correct that there is a bug in this part of the code. > > The full-length string from 'set missing' is tested against the > > start of the field contents (after removing leading whitespace); > > then the subsequent character is tested to see if it is whitespace > > rather than a continuation of whatever string is in the field. > > So it works with a tab-separated *.csv file because <tab> counts > > as whitespace, but fails with a comma-separated file because the > > comma is mis-interpreted as part of the field content. > > Thanks Ethan, > > First comment: a tab separated file is not a CSV file. The C mean comma > separated. In practice this is not true. Pretty much any program I know of that supports *.csv files allows you to specify what character is used as a field separator. Quoting Wikipedia: "the term "CSV" also denotes some closely related delimiter-separated formats that use different field delimiters. These include tab-separated values and space-separated values. A delimiter that is not present in the field data (such as tab) keeps the format parsing simple. These alternate delimiter-separated files are often even given a .csv extension, despite the use of a non-comma field separator." > I would suggest that the correct, structured way to do this is to break > into fields using the current field separator, then test whether any > fields match the missing string ( with the white-space caveats ). > > If I follow your explanation, it would seem that currently the whole > line is scanned for the 'missing' string before it is split into fields, > or it is being parsed twice. Not quite. The input line is scanned for field separators, the start of each field is noted, then it goes back to process them one-by-one. > > 2, 2, ignore A, 2 > 3, 3, ignore B, 3 > 4, 4, ignore, 4 > > > IMO 2 and 3 should not match since the field is not equal to the > 'missing' string but simply contains it. This sounds like asking for > trouble. Allowing white space seems sensible flexibility on insisting on > an exact match since it is often added for human readability, as is the > case here. > > Only something which IS the 'missing' string or the string with leading > and/or trailing white-space should match, IMO. The conventional indication of missing data in a *.csv file is simply an empty field. This obviously is not possible in a whitespace-separated file. Gnuplot's use of "set missing" is outside any standard practice I know of for csv files, so anything we choose is likely to strike someone as wrong. For instance, RFC-4180, the closest thing to a csv standard, states that "any field may be quoted with double quotes". So in the example above, should we ignore this line? 5, 5, "ignore", 5 This one? 5, 5, " ignore ", 5 Ethan > > Peter. |
|
From: Tatsuro M. <tma...@ya...> - 2016-06-14 22:20:46
|
The change 2016-06-14 Bastian Maerkisch <bma...@we...> * src/readline.c src/plot.c config/config.mgw config/mingw/Makefile: Allow console mode gnuplot on Windows to use GNU readline, which is more powerful than the builtin code, but still does not handle Unicode input on Windows. For GNU readline, we had discussed the below: http://gnuplot.10905.n7.nabble.com/readline-to-gnuplot-exe-console-mode-for-windows-td11173.html I at the moment do not link the GNU readline for my CVS binary distribution unless a new statement will appear about the GNU readline. Tatsuro |
|
From: <pl...@pi...> - 2016-06-14 22:00:30
|
On 14/06/16 21:00, Ethan A Merritt wrote: > On Tuesday, 14 June, 2016 13:10:13 pl...@pi... wrote: >> On 14/06/16 12:52, Allin Cottrell wrote: >>> On Tue, 14 Jun 2016, pl...@pi... wrote: >>> >>>> On 14/06/16 09:43, pl...@pi... wrote: >>>>> >>>>> I have a data csv datafile which uses -99.99 as it missing data value. >>>>> >>>>> with the following settings the 'missing' data are getting plotted. >>>>> >>>>> set datafile separator "," >>>>> set datafile missing "-99.99" >>>>> >>>>> >>>>> show datafile missing >>>>> >>>>> "-99.99" in datafile is interpreted as missing value >>>>> > > I think you are correct that there is a bug in this part of the code. > The full-length string from 'set missing' is tested against the > start of the field contents (after removing leading whitespace); > then the subsequent character is tested to see if it is whitespace > rather than a continuation of whatever string is in the field. > So it works with a tab-separated *.csv file because <tab> counts > as whitespace, but fails with a comma-separated file because the > comma is mis-interpreted as part of the field content. Thanks Ethan, First comment: a tab separated file is not a CSV file. The C mean comma separated. Your explanation seems to confirm my intuitive guess about how this was being processed. There should never be a question of the 'missing' test seeing the following comma since it is not the content of a field and I think that is the origin of the bug. I would suggest that the correct, structured way to do this is to break into fields using the current field separator, then test whether any fields match the missing string ( with the white-space caveats ). If I follow your explanation, it would seem that currently the whole line is scanned for the 'missing' string before it is split into fields, or it is being parsed twice. It seems logical that the line be split into fields before trying to test the value of any field for any condition. This appears not to be the case at the moment. > > This should be fixed. > The subsequent character should be tested for > <next character is either whitespace or field-separator>. > Or maybe it should be > <next character is field-separator (which might be whitespace)>. > I'm not sure which is correct. > > The difference would matter in a case like this: > > set datafile separator comma > set datafile missing "ignore" > > plot '-' using 1:3 > 1, 1, 1, 1 > 2, 2, ignore A, 2 > 3, 3, ignore B, 3 > 4, 4, ignore, 4 > 5, 5, 5, 5 > e > > In current gnuplot lines 2 and 3 will be treated as missing > but line 4 will not. > Should a fix result in only line 4 being ignored? > Or should all three lines be ignored? > >>>>> >>>>> As a wild guess I tried the following and the missing data now get >>>>> correctly removed. >>>>> >>>>> set datafile missing "-99.99," >>>>> >>>>> >>>>> This seems to be an illogical order of parsing. > > That will only work if there is whitespace following the comma. > So it's not a guaranteed work-around. > > Ethan > > >>>>> >>>>> Surely the data line needs to be parsed into its constituent data >>>>> columns before trying to detect the missing data string. >>>>> >>>>> Regards, Peter >>>>> >>>>> >>>>> gnuplot> show version >>>>> >>>>> G N U P L O T >>>>> Version 5.0 patchlevel 1 last modified 2015-06-07 >>>>> >>>>> Copyright (C) 1986-1993, 1998, 2004, 2007-2015 >>>>> Thomas Williams, Colin Kelley and many others >>>>> >>>>> gnuplot home: http://www.gnuplot.info >>>>> faq, bugs, etc: type "help FAQ" >>>>> immediate help: type "help" (plot window: hit 'h') >>>>> >>>>> >>>>> >>>> >>>> Just to complete this here is a sample line from the file displaying >>>> this bug. >>>> >>>> 1958, 06, 21351, 1958.4548, -99.99, -99.99, 317.25, 315.14, >>>> 317.25, 315.14 >>> >>> Doesn't it invite undefined behavior if you set "," as separator but >>> then also include spaces between the values? >>> >>> Allin Cottrell >>> >> >> I'm not including anything, I have some data provided that I need to >> plot with gnuplot. >> I've always found gnuplot smart enough to deal with most things that >> I've thrown at it. If this is not a bug I could always preprocess the >> data to remove the commas. >> >> The question remains as to whether this is a bug or not. >> >> The fourth field in that line is " -99.99" when using comma separator. >> >> This data is supplied with a missing VALUE of -99.99 . Will this match >> gnuplot's datafile missing defined as a string "-99.99" ? That depends >> upon how the equality test is done in a language that has fuzzy >> variable types. >> >> But that does not account for the behaviour of it working with datafile >> missing set to "-99.99," >> >> That seems to clearly indicate that there is logical problem here. The >> comma should no longer be there in the data field since it is the field >> separator. >> >> Peter. >> >> >> Thanks. 2, 2, ignore A, 2 3, 3, ignore B, 3 4, 4, ignore, 4 IMO 2 and 3 should not match since the field is not equal to the 'missing' string but simply contains it. This sounds like asking for trouble. Allowing white space seems sensible flexibility on insisting on an exact match since it is often added for human readability, as is the case here. Only something which IS the 'missing' string or the string with leading and/or trailing white-space should match, IMO. Peter. |
|
From: Ethan A M. <sf...@us...> - 2016-06-14 20:03:41
|
On Tuesday, 14 June, 2016 13:10:13 pl...@pi... wrote: > On 14/06/16 12:52, Allin Cottrell wrote: > > On Tue, 14 Jun 2016, pl...@pi... wrote: > > > >> On 14/06/16 09:43, pl...@pi... wrote: > >>> > >>> I have a data csv datafile which uses -99.99 as it missing data value. > >>> > >>> with the following settings the 'missing' data are getting plotted. > >>> > >>> set datafile separator "," > >>> set datafile missing "-99.99" > >>> > >>> > >>> show datafile missing > >>> > >>> "-99.99" in datafile is interpreted as missing value > >>> I think you are correct that there is a bug in this part of the code. The full-length string from 'set missing' is tested against the start of the field contents (after removing leading whitespace); then the subsequent character is tested to see if it is whitespace rather than a continuation of whatever string is in the field. So it works with a tab-separated *.csv file because <tab> counts as whitespace, but fails with a comma-separated file because the comma is mis-interpreted as part of the field content. This should be fixed. The subsequent character should be tested for <next character is either whitespace or field-separator>. Or maybe it should be <next character is field-separator (which might be whitespace)>. I'm not sure which is correct. The difference would matter in a case like this: set datafile separator comma set datafile missing "ignore" plot '-' using 1:3 1, 1, 1, 1 2, 2, ignore A, 2 3, 3, ignore B, 3 4, 4, ignore, 4 5, 5, 5, 5 e In current gnuplot lines 2 and 3 will be treated as missing but line 4 will not. Should a fix result in only line 4 being ignored? Or should all three lines be ignored? > >>> > >>> As a wild guess I tried the following and the missing data now get > >>> correctly removed. > >>> > >>> set datafile missing "-99.99," > >>> > >>> > >>> This seems to be an illogical order of parsing. That will only work if there is whitespace following the comma. So it's not a guaranteed work-around. Ethan > >>> > >>> Surely the data line needs to be parsed into its constituent data > >>> columns before trying to detect the missing data string. > >>> > >>> Regards, Peter > >>> > >>> > >>> gnuplot> show version > >>> > >>> G N U P L O T > >>> Version 5.0 patchlevel 1 last modified 2015-06-07 > >>> > >>> Copyright (C) 1986-1993, 1998, 2004, 2007-2015 > >>> Thomas Williams, Colin Kelley and many others > >>> > >>> gnuplot home: http://www.gnuplot.info > >>> faq, bugs, etc: type "help FAQ" > >>> immediate help: type "help" (plot window: hit 'h') > >>> > >>> > >>> > >> > >> Just to complete this here is a sample line from the file displaying > >> this bug. > >> > >> 1958, 06, 21351, 1958.4548, -99.99, -99.99, 317.25, 315.14, > >> 317.25, 315.14 > > > > Doesn't it invite undefined behavior if you set "," as separator but > > then also include spaces between the values? > > > > Allin Cottrell > > > > I'm not including anything, I have some data provided that I need to > plot with gnuplot. > I've always found gnuplot smart enough to deal with most things that > I've thrown at it. If this is not a bug I could always preprocess the > data to remove the commas. > > The question remains as to whether this is a bug or not. > > The fourth field in that line is " -99.99" when using comma separator. > > This data is supplied with a missing VALUE of -99.99 . Will this match > gnuplot's datafile missing defined as a string "-99.99" ? That depends > upon how the equality test is done in a language that has fuzzy > variable types. > > But that does not account for the behaviour of it working with datafile > missing set to "-99.99," > > That seems to clearly indicate that there is logical problem here. The > comma should no longer be there in the data field since it is the field > separator. > > Peter. > > > > > ------------------------------------------------------------------------------ > What NetFlow Analyzer can do for you? Monitors network bandwidth and traffic > patterns at an interface-level. Reveals which users, apps, and protocols are > consuming the most bandwidth. Provides multi-vendor support for NetFlow, > J-Flow, sFlow and other flows. Make informed decisions using capacity > planning reports. https://ad.doubleclick.net/ddm/clk/305295220;132659582;e > _______________________________________________ > gnuplot-beta mailing list > gnu...@li... > Membership management via: https://lists.sourceforge.net/lists/listinfo/gnuplot-beta |
|
From: <pl...@pi...> - 2016-06-14 16:25:21
|
On 14/06/16 12:52, Allin Cottrell wrote: > On Tue, 14 Jun 2016, pl...@pi... wrote: > >> On 14/06/16 09:43, pl...@pi... wrote: >>> >>> I have a data csv datafile which uses -99.99 as it missing data value. >>> >>> with the following settings the 'missing' data are getting plotted. >>> >>> set datafile separator "," >>> set datafile missing "-99.99" >>> >>> >>> show datafile missing >>> >>> "-99.99" in datafile is interpreted as missing value >>> >>> >>> As a wild guess I tried the following and the missing data now get >>> correctly removed. >>> >>> set datafile missing "-99.99," >>> >>> >>> This seems to be an illogical order of parsing. >>> >>> Surely the data line needs to be parsed into its constituent data >>> columns before trying to detect the missing data string. >>> >>> Regards, Peter >>> >>> >>> gnuplot> show version >>> >>> G N U P L O T >>> Version 5.0 patchlevel 1 last modified 2015-06-07 >>> >>> Copyright (C) 1986-1993, 1998, 2004, 2007-2015 >>> Thomas Williams, Colin Kelley and many others >>> >>> gnuplot home: http://www.gnuplot.info >>> faq, bugs, etc: type "help FAQ" >>> immediate help: type "help" (plot window: hit 'h') >>> >>> >>> >> >> Just to complete this here is a sample line from the file displaying >> this bug. >> >> 1958, 06, 21351, 1958.4548, -99.99, -99.99, 317.25, 315.14, >> 317.25, 315.14 > > Doesn't it invite undefined behavior if you set "," as separator but > then also include spaces between the values? > > Allin Cottrell > I'm not including anything, I have some data provided that I need to plot with gnuplot. I've always found gnuplot smart enough to deal with most things that I've thrown at it. If this is not a bug I could always preprocess the data to remove the commas. The question remains as to whether this is a bug or not. The fourth field in that line is " -99.99" when using comma separator. This data is supplied with a missing VALUE of -99.99 . Will this match gnuplot's datafile missing defined as a string "-99.99" ? That depends upon how the equality test is done in a language that has fuzzy variable types. But that does not account for the behaviour of it working with datafile missing set to "-99.99," That seems to clearly indicate that there is logical problem here. The comma should no longer be there in the data field since it is the field separator. Peter. |
|
From: Allin C. <cot...@wf...> - 2016-06-14 12:16:29
|
On Tue, 14 Jun 2016, pl...@pi... wrote: > On 14/06/16 09:43, pl...@pi... wrote: >> >> I have a data csv datafile which uses -99.99 as it missing data value. >> >> with the following settings the 'missing' data are getting plotted. >> >> set datafile separator "," >> set datafile missing "-99.99" >> >> >> show datafile missing >> >> "-99.99" in datafile is interpreted as missing value >> >> >> As a wild guess I tried the following and the missing data now get >> correctly removed. >> >> set datafile missing "-99.99," >> >> >> This seems to be an illogical order of parsing. >> >> Surely the data line needs to be parsed into its constituent data >> columns before trying to detect the missing data string. >> >> Regards, Peter >> >> >> gnuplot> show version >> >> G N U P L O T >> Version 5.0 patchlevel 1 last modified 2015-06-07 >> >> Copyright (C) 1986-1993, 1998, 2004, 2007-2015 >> Thomas Williams, Colin Kelley and many others >> >> gnuplot home: http://www.gnuplot.info >> faq, bugs, etc: type "help FAQ" >> immediate help: type "help" (plot window: hit 'h') >> >> >> > > Just to complete this here is a sample line from the file displaying > this bug. > > 1958, 06, 21351, 1958.4548, -99.99, -99.99, 317.25, 315.14, > 317.25, 315.14 Doesn't it invite undefined behavior if you set "," as separator but then also include spaces between the values? Allin Cottrell |
|
From: <pl...@pi...> - 2016-06-14 11:25:26
|
On 14/06/16 09:43, pl...@pi... wrote: > > I have a data csv datafile which uses -99.99 as it missing data value. > > with the following settings the 'missing' data are getting plotted. > > set datafile separator "," > set datafile missing "-99.99" > > > show datafile missing > > "-99.99" in datafile is interpreted as missing value > > > As a wild guess I tried the following and the missing data now get > correctly removed. > > set datafile missing "-99.99," > > > This seems to be an illogical order of parsing. > > Surely the data line needs to be parsed into its constituent data > columns before trying to detect the missing data string. > > Regards, Peter > > > gnuplot> show version > > G N U P L O T > Version 5.0 patchlevel 1 last modified 2015-06-07 > > Copyright (C) 1986-1993, 1998, 2004, 2007-2015 > Thomas Williams, Colin Kelley and many others > > gnuplot home: http://www.gnuplot.info > faq, bugs, etc: type "help FAQ" > immediate help: type "help" (plot window: hit 'h') > > > Just to complete this here is a sample line from the file displaying this bug. 1958, 06, 21351, 1958.4548, -99.99, -99.99, 317.25, 315.14, 317.25, 315.14 It seems that the data is being parsed into fields using the default space delimiter when looking for the "missing" string. It should be using proper user-defined delimiter. Peter. |
|
From: <pl...@pi...> - 2016-06-14 10:00:45
|
I have a data csv datafile which uses -99.99 as it missing data value.
with the following settings the 'missing' data are getting plotted.
set datafile separator ","
set datafile missing "-99.99"
show datafile missing
"-99.99" in datafile is interpreted as missing value
As a wild guess I tried the following and the missing data now get
correctly removed.
set datafile missing "-99.99,"
This seems to be an illogical order of parsing.
Surely the data line needs to be parsed into its constituent data
columns before trying to detect the missing data string.
Regards, Peter
gnuplot> show version
G N U P L O T
Version 5.0 patchlevel 1 last modified 2015-06-07
Copyright (C) 1986-1993, 1998, 2004, 2007-2015
Thomas Williams, Colin Kelley and many others
gnuplot home: http://www.gnuplot.info
faq, bugs, etc: type "help FAQ"
immediate help: type "help" (plot window: hit 'h')
|
|
From: Daniel J S. <dan...@ie...> - 2016-06-08 21:26:04
|
On 06/08/2016 03:20 PM, Ethan A Merritt wrote: > On Wednesday, 08 June, 2016 15:05:16 Daniel J Sebald wrote: >> On 06/08/2016 01:39 PM, Ethan A Merritt wrote: > >>>> piecewise.dem (1) Piecewise function sampling >>>> matrix_index.dem (1) Data file contains labeled ascii matrices >>>> [annotation off screen] >>> >>> These demos did not exist prior to version 5, so I'm not sure what >>> you are comparing to. The output looks normal when tested here. >> >> I'm seeing the y-axis text right up against the screen edge, and on the >> right side is three characters of white space. But, I see that margins >> are specified in this example. Note, however, that I'm seeing an error >> message for the "set multiplot" line: >> >> gnuplot> load 'matrix_index.dem' >> "matrix_index.dem", line 29: warning: must give margins and >> spacing, continue with spacing of 0.05 >> Hit return to continue > > That warning seems unnecessarily alarming. > All it means is that in the absence of an explicit spacing keyword > it will use the default value. I agree. gnuplot does defaults all the time. I would say though the default is kind of small and I'd prefer it to be in characters because that is what "set tmargin", etc. uses. Screen is good for fine-tuning, but something near whole numbers, rather than .05, is a better place to start. >> If I just remove the "margins etc" option, the warning disappears and >> the spacing looks good, very balanced. Should we just remove the >> margins? Or keep them in as a way of testing the option? Here's what >> looks good for me: > > [shrug] Tastes differ. > I think it looks better with identical left and right margins. > To me the default looks lopsided. That's because it is, now that I look at. The top two plots have titles so they are smaller than the bottom two plots. Setting the margins must fix the axes locations rather than the overall space. Option margins char 5,3,2,4 spacing char 6,3 looks much better than the current setting though. Dan |
|
From: Ethan A M. <sf...@us...> - 2016-06-08 20:24:12
|
On Wednesday, 08 June, 2016 15:05:16 Daniel J Sebald wrote: > On 06/08/2016 01:39 PM, Ethan A Merritt wrote: > >> piecewise.dem (1) Piecewise function sampling > >> matrix_index.dem (1) Data file contains labeled ascii matrices > >> [annotation off screen] > > > > These demos did not exist prior to version 5, so I'm not sure what > > you are comparing to. The output looks normal when tested here. > > I'm seeing the y-axis text right up against the screen edge, and on the > right side is three characters of white space. But, I see that margins > are specified in this example. Note, however, that I'm seeing an error > message for the "set multiplot" line: > > gnuplot> load 'matrix_index.dem' > "matrix_index.dem", line 29: warning: must give margins and > spacing, continue with spacing of 0.05 > Hit return to continue That warning seems unnecessarily alarming. All it means is that in the absence of an explicit spacing keyword it will use the default value. > If I just remove the "margins etc" option, the warning disappears and > the spacing looks good, very balanced. Should we just remove the > margins? Or keep them in as a way of testing the option? Here's what > looks good for me: [shrug] Tastes differ. I think it looks better with identical left and right margins. To me the default looks lopsided. Ethan |
|
From: Daniel J S. <dan...@ie...> - 2016-06-08 20:05:54
|
On 06/08/2016 01:39 PM, Ethan A Merritt wrote: > On Wednesday, 08 June, 2016 11:07:28 Daniel J Sebald wrote: >> I've been stepping through the demos, and my general observation from >> what I remember is that there is now much less white space around plots. >> The behavior is the same in two different terminals (x11 and qt). >> I've gone back to about -D2016-05-25 in the repository to see if this is >> a current inadvertent change. It's the same as current development version. > > The default margins for "set view map; splot ..." were changed to be > more like those for an equivalent 2D plot. So that category of plot > is expected to have less white space around it than in old gnuplot > versions. OK, so it looks like primarily the issue is that a lot of demos need some fine tuning...or remove fine tuning that was geared toward older layout. >> Has anyone else noticed this? Perhaps it was intentional, and in some >> sense I don't mind it because it makes the plots use more of the >> available window screen. I mean, typically when I create a plot for a >> publication I will reduce the whitespace to minimal because word >> processors can always add it back in. That can be done changing the >> margins or post-gnuplot with some cropping tool specific to, say, eps or >> something. >> >> However, as far as the demos, the plots don't feel as though they have a >> good white space balance for several plots. Here are some notable ones: > > >> fillcrvs.dem (7) world.dat plotted with filledcurves >> multiplt.dem (1) > > I have no explanation for any changes seen in these Ah, there's a margin command in that example that was probably for tweaking the older layout behavior. This change looks much better: Index: gnuplot/demo/fillcrvs.dem =================================================================== RCS file: /cvsroot/gnuplot/gnuplot/demo/fillcrvs.dem,v retrieving revision 1.6 diff -u -r1.6 fillcrvs.dem --- gnuplot/demo/fillcrvs.dem 19 May 2007 20:35:53 -0000 1.6 +++ gnuplot/demo/fillcrvs.dem 8 Jun 2016 20:00:20 -0000 @@ -70,7 +70,6 @@ set object 1 rect from graph 0, 0 to graph 1, 1 behind fc rgb "#afffff" fillstyle solid 1.00 border -1 set xrange [ -180.000 : 180.000 ] set yrange [ -70.0000 : 80.0000 ] -set lmargin 1 plot 'world.dat' with filledcurve notitle fs solid 1.0 lc rgb 'dark-goldenrod' pause -1 'Press Return to continue' >> [The above two illustrate how plot are is slightly skewed to the left] >> heatmaps.dem (1) (2) (3) (4) > > These are set view map + splot, so changes are expected OK, map view, yeah there's quite a few of those. The historic map size was kind of small. This is the change that was noted in this bug report: https://sourceforge.net/p/gnuplot/bugs/1802/ I will follow up on that one on Sourceforge. >> piecewise.dem (1) Piecewise function sampling >> matrix_index.dem (1) Data file contains labeled ascii matrices >> [annotation off screen] > > These demos did not exist prior to version 5, so I'm not sure what > you are comparing to. The output looks normal when tested here. I'm seeing the y-axis text right up against the screen edge, and on the right side is three characters of white space. But, I see that margins are specified in this example. Note, however, that I'm seeing an error message for the "set multiplot" line: gnuplot> load 'matrix_index.dem' "matrix_index.dem", line 29: warning: must give margins and spacing, continue with spacing of 0.05 Hit return to continue If I just remove the "margins etc" option, the warning disappears and the spacing looks good, very balanced. Should we just remove the margins? Or keep them in as a way of testing the option? Here's what looks good for me: Index: gnuplot/demo/matrix_index.dem =================================================================== RCS file: /cvsroot/gnuplot/gnuplot/demo/matrix_index.dem,v retrieving revision 1.2 diff -u -r1.2 matrix_index.dem --- gnuplot/demo/matrix_index.dem 13 Jan 2015 18:17:29 -0000 1.2 +++ gnuplot/demo/matrix_index.dem 8 Jun 2016 20:00:20 -0000 @@ -26,7 +26,7 @@ set xrange [] noextend set yrange [] noextend -set multiplot layout 2,2 margins char 3,3,2,4 title "{/:Bold Data file contains labeled ascii matrices} +set multiplot layout 2,2 margins char 5,3,2,4 spacing char 6,3 title "{/:Bold Data file contains labeled ascii matrices} set title "Y range should be the same" plot '$MATRICES' nonuniform matrix i "set3" w image title "index 'set3'" Dan |