You can subscribe to this list here.
| 2001 |
Jan
|
Feb
(1) |
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
|
Sep
|
Oct
|
Nov
|
Dec
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2002 |
Jan
(1) |
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
|
Nov
(1) |
Dec
|
| 2003 |
Jan
|
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
(83) |
Nov
(57) |
Dec
(111) |
| 2004 |
Jan
(38) |
Feb
(121) |
Mar
(107) |
Apr
(241) |
May
(102) |
Jun
(190) |
Jul
(239) |
Aug
(158) |
Sep
(184) |
Oct
(193) |
Nov
(47) |
Dec
(68) |
| 2005 |
Jan
(190) |
Feb
(105) |
Mar
(99) |
Apr
(65) |
May
(92) |
Jun
(250) |
Jul
(197) |
Aug
(128) |
Sep
(101) |
Oct
(183) |
Nov
(186) |
Dec
(42) |
| 2006 |
Jan
(102) |
Feb
(122) |
Mar
(154) |
Apr
(196) |
May
(181) |
Jun
(281) |
Jul
(310) |
Aug
(198) |
Sep
(145) |
Oct
(188) |
Nov
(134) |
Dec
(90) |
| 2007 |
Jan
(134) |
Feb
(181) |
Mar
(157) |
Apr
(57) |
May
(81) |
Jun
(204) |
Jul
(60) |
Aug
(37) |
Sep
(17) |
Oct
(90) |
Nov
(122) |
Dec
(72) |
| 2008 |
Jan
(130) |
Feb
(108) |
Mar
(160) |
Apr
(38) |
May
(83) |
Jun
(42) |
Jul
(75) |
Aug
(16) |
Sep
(71) |
Oct
(57) |
Nov
(59) |
Dec
(152) |
| 2009 |
Jan
(73) |
Feb
(213) |
Mar
(67) |
Apr
(40) |
May
(46) |
Jun
(82) |
Jul
(73) |
Aug
(57) |
Sep
(108) |
Oct
(36) |
Nov
(153) |
Dec
(77) |
| 2010 |
Jan
(42) |
Feb
(171) |
Mar
(150) |
Apr
(6) |
May
(22) |
Jun
(34) |
Jul
(31) |
Aug
(38) |
Sep
(32) |
Oct
(59) |
Nov
(13) |
Dec
(62) |
| 2011 |
Jan
(114) |
Feb
(139) |
Mar
(126) |
Apr
(51) |
May
(53) |
Jun
(29) |
Jul
(41) |
Aug
(29) |
Sep
(35) |
Oct
(87) |
Nov
(42) |
Dec
(20) |
| 2012 |
Jan
(111) |
Feb
(66) |
Mar
(35) |
Apr
(59) |
May
(71) |
Jun
(32) |
Jul
(11) |
Aug
(48) |
Sep
(60) |
Oct
(87) |
Nov
(16) |
Dec
(38) |
| 2013 |
Jan
(5) |
Feb
(19) |
Mar
(41) |
Apr
(47) |
May
(14) |
Jun
(32) |
Jul
(18) |
Aug
(68) |
Sep
(9) |
Oct
(42) |
Nov
(12) |
Dec
(10) |
| 2014 |
Jan
(14) |
Feb
(139) |
Mar
(137) |
Apr
(66) |
May
(72) |
Jun
(142) |
Jul
(70) |
Aug
(31) |
Sep
(39) |
Oct
(98) |
Nov
(133) |
Dec
(44) |
| 2015 |
Jan
(70) |
Feb
(27) |
Mar
(36) |
Apr
(11) |
May
(15) |
Jun
(70) |
Jul
(30) |
Aug
(63) |
Sep
(18) |
Oct
(15) |
Nov
(42) |
Dec
(29) |
| 2016 |
Jan
(37) |
Feb
(48) |
Mar
(59) |
Apr
(28) |
May
(30) |
Jun
(43) |
Jul
(47) |
Aug
(14) |
Sep
(21) |
Oct
(26) |
Nov
(10) |
Dec
(2) |
| 2017 |
Jan
(26) |
Feb
(27) |
Mar
(44) |
Apr
(11) |
May
(32) |
Jun
(28) |
Jul
(75) |
Aug
(45) |
Sep
(35) |
Oct
(285) |
Nov
(99) |
Dec
(16) |
| 2018 |
Jan
(8) |
Feb
(8) |
Mar
(42) |
Apr
(35) |
May
(23) |
Jun
(12) |
Jul
(16) |
Aug
(11) |
Sep
(8) |
Oct
(16) |
Nov
(5) |
Dec
(8) |
| 2019 |
Jan
(9) |
Feb
(28) |
Mar
(4) |
Apr
(10) |
May
(7) |
Jun
(4) |
Jul
(4) |
Aug
|
Sep
(4) |
Oct
|
Nov
(23) |
Dec
(3) |
| 2020 |
Jan
(19) |
Feb
(3) |
Mar
(22) |
Apr
(17) |
May
(10) |
Jun
(69) |
Jul
(18) |
Aug
(23) |
Sep
(25) |
Oct
(11) |
Nov
(20) |
Dec
(9) |
| 2021 |
Jan
(1) |
Feb
(7) |
Mar
(9) |
Apr
|
May
(1) |
Jun
(8) |
Jul
(6) |
Aug
(8) |
Sep
(7) |
Oct
|
Nov
(2) |
Dec
(23) |
| 2022 |
Jan
(23) |
Feb
(9) |
Mar
(9) |
Apr
|
May
(8) |
Jun
(1) |
Jul
(6) |
Aug
(8) |
Sep
(30) |
Oct
(5) |
Nov
(4) |
Dec
(6) |
| 2023 |
Jan
(2) |
Feb
(5) |
Mar
(7) |
Apr
(3) |
May
(8) |
Jun
(45) |
Jul
(8) |
Aug
|
Sep
(2) |
Oct
(14) |
Nov
(7) |
Dec
(2) |
| 2024 |
Jan
(4) |
Feb
(4) |
Mar
|
Apr
(7) |
May
(2) |
Jun
(1) |
Jul
|
Aug
(5) |
Sep
|
Oct
|
Nov
(4) |
Dec
(14) |
| 2025 |
Jan
(22) |
Feb
(6) |
Mar
(5) |
Apr
(14) |
May
(6) |
Jun
(11) |
Jul
(19) |
Aug
|
Sep
(17) |
Oct
(1) |
Nov
(2) |
Dec
(18) |
| 2026 |
Jan
|
Feb
|
Mar
(5) |
Apr
|
May
(2) |
Jun
(1) |
Jul
(6) |
Aug
(1) |
Sep
|
Oct
|
Nov
|
Dec
|
|
From: <pl...@pi...> - 2007-10-25 00:44:57
|
On Wed, 24 Oct 2007 19:20:12 +0200, Ethan Merritt <merritt@u.washington.edu> wrote: > On Wednesday 24 October 2007 05:45, pl...@pi... wrote: >> >> if you want to post that as an example it should probably by plotting >> against ($1+old_x)/2 to prevent a shift in the position of the >> derivative. >> >> BTW the running mean example also suffers from this defect as is clearly >> seen in the plot that Ethan linked to. The running mean is offset 2.5 >> data >> points to the right. > > The running mean example is correct for the typical use case. > Have a look in your local paper for the stock price reports, > or the weather summaries, etc. > > In calculating a running average you typically do not known the value of > future points, only of the points measured to date. > So you necessarily average over the previous 5 points, as stated in the > plot title. > > This is not the same as for local approximation of the derivative. > Yes the title does clearly state what is shown. I did not say it was incorrect, I said it was a defect. Hand-wavy discussions about some "typical usage" are not much help in science and I certainly would not look to my local paper for an reference on the correct way to process data. I posted that comment since it is a common pitfall in calculating a running mean which also applied to the derviative code snip. Please dont take it the wrong way. ;) |
|
From: <pl...@pi...> - 2007-10-25 00:28:18
|
On Tue, 23 Oct 2007 22:01:33 +0200, <ln...@dr...> wrote: >> >> Yes one of the problems with splines is that the contraints placed on >> continuity and the fact that it must pass through all data points can >> lead >> to some fairly unexpected and sometimes extreme excursions from the zone >> where the real data lie. >> > I beg to disagree: you certainly are not forced to go through all your > data > points for a smoothed spline. Of course, csplines do this way, but > acsplines not > (and bezier curves also do not necessarily pass through all the points). You are correct to an extent but acsplines are still basically applying a cspline and hence the artificial constraints that are the basis of that technique. Massaging the wieghts to get it to "look better" is adding error to error it does not correct the fundemental problem. For example csplines have problems following rapid changes, ie they dont follow and then they tend to overshoot badly. If you use acsplines and play with the wieghts around the change you may play down the overshoot by reducing the surrounding wieghts. But in doing this you mask the change that is real. You are further corrupting the data and masking real effects, not improving the fit. The ability to add a column of weights is valuable but these values should have some experimental justification not just be used as arbitary fiddle factors. I dont think this approach addresses my critisim of csplines. It may be that you are trying to clobber some fairly high frequency experimental noise on an effect that is fairly smooth in nature. In this case, as you point out, fiddling coeffs is a bit like a hand drawn line and taking a tangent with a ruler and pencil. If that fits your needs you're all set, it's good that you were able to find a solution for the derivatives. I think this case in point highlights a lack in gnuplot's offering in the "smoothing" department. I've touched on this before. csplines are pretty heavy handed when it comes to rounding things off and produce some severe artifacts unless the underlying nature of the data is itself long smooth curves. I know Hans-Bernhard is going to cry foul and say that gnuplot should not be doing data processing. The fact is it does. The trouble is cspines are a bit special and the current offering of csplines or bezier is a bit one-sided. It would be nice to have an alternative that suits other sorts of data. Some sort of mean or wieghted mean for example. regards, Peter. |
|
From: Philipp K. J. <ja...@ie...> - 2007-10-24 23:02:21
|
On Wednesday 24 October 2007 12:10, you wrote: > Philipp K. Janert wrote: > > Now, the real question is, what are the relative > > advantages for different choices of w(d)? > > Well, the primary advantage of the current choice is that it _is_ the > current choice. dgrid3d must have been in gnuplot for over a decade > now. Changing it now would require careful consideration to ensure > backward compatibility. That's a good point. What kind of things do are you thinking of that depend on the current implementation? Changing the weight function does not seem to break any existing code, as far as I can see. The best way forward might be to put the choice into the hands of the USER anyway. We can add an option to dgrid3d which selects the filter, defaulting to the current one (this would guarantee that the behaviour of existing gnuplot scripts does not change. > > > What > > I am missing most strongly in the current > > implementation is a way to control the radius > > over which data points are averaged to form > > the smooth approximation - control over the > > "band width" of the low pass filter. > > Well, as I said before, there is no such radius, thus nothing to be > controlled. OTOH, that's how the algorithm manages to work without > knowing anything about the data. Yes, that's true. But I think that's exactly what would be nice to have - to put the ability to control the range into the hands of users, who DO know their data! For example, if I know that my data is very smooth, but not on a grid, I might want to use dgrid3d with a small averaging range, just to get a surface drawn. But when I know my data to be noisy, I might want to have a wide averaging range, to get some of the noise out. The point is that putting control into the hands of the user allows the user to get more value out of the tool. (We can still provide the current behaviour as default, so there is no change unless requested.) > I.e. the algorithm is invariant under > some rescaling of the input. The surface generated by dgrid3d for > > splot 'data' using ($1*s_xy):($2*s_xy):($3*s_z) > > looks the same for any s_xy and s_z (setting aside different choices by > the autoscaling algorithm). Only if you scale x and y differently, the > result changes. One aspect where this helps is that you don't have to > adjust the settings for every plot. > > Maybe it's the wording in the documentation that's misleading here. > dgrid3d is a kind of low pass filter, but it's a rather different kind > than you're used to. It's sort of an auto-adjusting low-pass filter. > It has no predefined limiting frequency. |
|
From: Ethan M. <merritt@u.washington.edu> - 2007-10-24 19:20:44
|
On Wednesday 24 October 2007 06:16, Fr=C3=A9d=C3=A9ric Gava wrote: > Dear Gnuplot developpers and Ethan Merritt, >=20 > I attach the complete files and outputs with the problem. I used gnuplot= =20 > 4.2 with Ubuntu 6.10. You can see in output.dvi that the key is in the=20 > middle The command file you give uses the command set key left plot 'test.dat' using 1:2 title "To be on the left" with linespoints I think if you turn on the key around the box you will see exactly what is happening. The key *is* drawn to the left, but the contents are right-justified. Whay you really want is: set key left Left box or maybe even set key left Left reverse box You can omit the "box" part after you understand what is going on. Yes, there are a lot of options to the "set key" command. This is the penalty for trying to give complete control to the user. =2D-=20 Ethan A Merritt |
|
From: <HBB...@t-...> - 2007-10-24 19:10:12
|
Philipp K. Janert wrote: > Now, the real question is, what are the relative > advantages for different choices of w(d)? Well, the primary advantage of the current choice is that it _is_ the current choice. dgrid3d must have been in gnuplot for over a decade now. Changing it now would require careful consideration to ensure backward compatibility. > What > I am missing most strongly in the current > implementation is a way to control the radius > over which data points are averaged to form > the smooth approximation - control over the > "band width" of the low pass filter. Well, as I said before, there is no such radius, thus nothing to be controlled. OTOH, that's how the algorithm manages to work without knowing anything about the data. I.e. the algorithm is invariant under some rescaling of the input. The surface generated by dgrid3d for splot 'data' using ($1*s_xy):($2*s_xy):($3*s_z) looks the same for any s_xy and s_z (setting aside different choices by the autoscaling algorithm). Only if you scale x and y differently, the result changes. One aspect where this helps is that you don't have to adjust the settings for every plot. Maybe it's the wording in the documentation that's misleading here. dgrid3d is a kind of low pass filter, but it's a rather different kind than you're used to. It's sort of an auto-adjusting low-pass filter. It has no predefined limiting frequency. |
|
From: Ethan M. <merritt@u.washington.edu> - 2007-10-24 18:16:31
|
On Wednesday 24 October 2007 05:45, pl...@pi... wrote: > > if you want to post that as an example it should probably by plotting > against ($1+old_x)/2 to prevent a shift in the position of the derivative. > > BTW the running mean example also suffers from this defect as is clearly > seen in the plot that Ethan linked to. The running mean is offset 2.5 data > points to the right. The running mean example is correct for the typical use case. Have a look in your local paper for the stock price reports, or the weather summaries, etc. In calculating a running average you typically do not known the value of future points, only of the points measured to date. So you necessarily average over the previous 5 points, as stated in the plot title. This is not the same as for local approximation of the derivative. -- Ethan A Merritt |
|
From: Philipp K. J. <ja...@ie...> - 2007-10-24 15:51:01
|
snip > > As you said the gaussian would be an obvious choice , having a finite > weight close to zero seems essencial. > snip > > Nice plots , I think you illustrate the point very well . There are some > obvious artifacts in the current model particularly heavy in the lower > order plots. > > I think your suggestion is a marked improvement from a theoretical point > of view and that your example points out that the magnitude of the > improvements possible by applying a more suitable weighting function. > > Nice contribution. I hope it gets adopted. Thanks. The next step up would be to have two-different scale factors as part of dgrid3d - one for the x- and y-directions, respectively. Then the weight function becomes: exp( - (x/x0)**2 + (y/y0)**2 ) which means that you can now control the radius of influence for both x and y IN THE UNITS OF THE DATA ITSELF, without having to resort to rescaling your data outside of gnuplot. I think this would be really useful. Best, Ph. |
|
From: Philipp K. J. <ja...@ie...> - 2007-10-24 15:36:55
|
Hans-Bernhard: I see. This explains our miscommunication: we use different terminology. I call w(d) the weight function (which is entirely independent of the data set), whereas you use this term for the fully normalized coefficient: w(d_i) / Sum_i w(d_i) Ok. Now, the real question is, what are the relative advantages for different choices of w(d)? What I am missing most strongly in the current implementation is a way to control the radius over which data points are averaged to form the smooth approximation - control over the "band width" of the low pass filter. With the suggestion I made, this can be done - without forcing the user to pre-process the data outside of gnuplot. It seems this would be a useful feature. Best, Ph. On Tuesday 23 October 2007 12:19, you wrote: > Philipp K. Janert wrote: > > I still don't think we are really on the same > > page about this. > > > > We agree that dgrid3d calculates a weighted > > average: > > z = Sum_i w_i(d) z_i / Sum_i w_i(d) > > where > > z_i is the i-th raw data point > > w_i(d) is the weight function > > Let's make that w(d_i). It's the distance d that depends on i, not the > function that turns a distance into a w. > > We at least partly disagree at this point. The actual weight function > is not w(d_i), but > > W_i := w(d_i) / Sum_j w(d_j) > > This is smooth (continuous even at d_i = 0), and it's limited to [0:1] > as weights should be. > > > The current implementation uses > > w(d) = 1/d**p if d > 0 > > w(d) = 1 if d == 0 > > No, it doesn't. It's not really w(d), but the final W_i that's set to > 1. This is achieved indirectly, by breaking out of the summation loop, > such that the result becomes the first input point with d==0. > > > The lesser reason is this: > > For d small, we multiply values by a large coefficient, > > form the sum, then divide the large coefficient > > out again. Numerically this is not very nice, because > > we loose significant digits if we multiply with something, > > sum, and then divide it out again. > > Multiplication doesn't lose digits at all, and summation only does where > it has to, irrespective of scale. The multiply/divide pair has > effectively no effect on numerical accuracy at all. In other words, the > given algorithm is quite unlikely to have any bigger problem with > round-off than a gaussian filter kernel would. |
|
From: <ga...@un...> - 2007-10-24 13:16:30
|
Dear Gnuplot developpers and Ethan Merritt, I attach the complete files and outputs with the problem. I used gnuplot 4.2 with Ubuntu 6.10. You can see in output.dvi that the key is in the middle test.dat : my data test.plt : the gnuplot commands test.tex : the tex output test.eps : the eps file needed by test.tex output.tex : a tex file to print the output output.dvi : the output Note that the problem is greater when I resize the output using the gnuplot command "set size 0.7,0.9" Best regards, Frédéric Gava |
|
From: Levente <ln...@dr...> - 2007-10-24 13:00:56
|
On Wed, 2007-10-24 at 14:45 +0200, pl...@pi... wrote: > On Wed, 24 Oct 2007 13:43:22 +0200, Levente Nov=C3=A1k =20 > <ln...@dr...> wrote: > > First of all, thanks to everybody who helped me! I succeeded in > > approximating the differential of data files the following way, using > > J=C3=BCrgen's example: > > > > # Approximation of the first derivative of x-y data > > # by the method: (y(n)-y(n-1))/(x(n)-x(n-1)) > > > > old_x =3D NaN > > old_y =3D NaN > > plot 'something.dat' using 1:(dx=3D$1-old_x,dy=3D$2-old_y,old_x=3D$1,ol= d_y=3D > > $2,dy/dx) > > > > Just to let know to whom it may be useful. Maybe a small demo file coul= d > > be written to demonstrate its use (I am pleased to post such a demo if > > this is OK). > > > Glad you found this feature useful. >=20 > It's quite new so it needs some different uses to see if does all that's = =20 > needed, thanks. >=20 > if you want to post that as an example it should probably by plotting =20 > against ($1+old_x)/2 to prevent a shift in the position of the derivative= . >=20 You are absolutely right, the x values need to be corrected (only I had so many datapoints that this shift was very small in my initial example). I will adapt the formula accordingly. Levente |
|
From: <pl...@pi...> - 2007-10-24 12:44:14
|
On Wed, 24 Oct 2007 13:43:22 +0200, Levente Novák <ln...@dr...> wrote: > On Mon, 2007-10-22 at 13:51 +0200, Juergen Wieferink wrote: > >> There is a new assignment operator (and a comma operator), so you >> can plot something like: >> >> old_x = NaN >> plot 'test.dat' using 1:(dx=$1-old_x,old_x=$1,dx) >> > > First of all, thanks to everybody who helped me! I succeeded in > approximating the differential of data files the following way, using > Jürgen's example: > > # Approximation of the first derivative of x-y data > # by the method: (y(n)-y(n-1))/(x(n)-x(n-1)) > > old_x = NaN > old_y = NaN > plot 'something.dat' using 1:(dx=$1-old_x,dy=$2-old_y,old_x=$1,old_y= > $2,dy/dx) > > Just to let know to whom it may be useful. Maybe a small demo file could > be written to demonstrate its use (I am pleased to post such a demo if > this is OK). > > Thanks again, folks! > > Levente > Glad you found this feature useful. It's quite new so it needs some different uses to see if does all that's needed, thanks. if you want to post that as an example it should probably by plotting against ($1+old_x)/2 to prevent a shift in the position of the derivative. BTW the running mean example also suffers from this defect as is clearly seen in the plot that Ethan linked to. The running mean is offset 2.5 data points to the right. regards. |
|
From: Levente <ln...@dr...> - 2007-10-24 11:44:52
|
On Mon, 2007-10-22 at 13:51 +0200, Juergen Wieferink wrote: > There is a new assignment operator (and a comma operator), so you > can plot something like: >=20 > old_x =3D NaN > plot 'test.dat' using 1:(dx=3D$1-old_x,old_x=3D$1,dx) >=20 First of all, thanks to everybody who helped me! I succeeded in approximating the differential of data files the following way, using J=C3=BCrgen's example: # Approximation of the first derivative of x-y data # by the method: (y(n)-y(n-1))/(x(n)-x(n-1)) old_x =3D NaN old_y =3D NaN plot 'something.dat' using 1:(dx=3D$1-old_x,dy=3D$2-old_y,old_x=3D$1,old_y= =3D $2,dy/dx) Just to let know to whom it may be useful. Maybe a small demo file could be written to demonstrate its use (I am pleased to post such a demo if this is OK). Thanks again, folks! Levente |
|
From: <ln...@dr...> - 2007-10-23 20:33:15
|
> On Tuesday 23 October 2007 13:01, ln...@dr... wrote: >> >> My biggest problem for the moment is the following: I did not find any= thing >> useful about 'assignment' in the help file (I have the most recent CVS= version >> of gnuplot, of course). How should I use it then? Is there a configure= -time >> switch needed to compile it into gnuplot? > > No, but it is a very recent feature. > > Peter <pl...@pi...> referred to it as an "assignment feature" b= ecause > the original trial implementation was in the form of a separate functio= n > named assign(). In the end, the same functionality was incorporated by > adding support for two standard operators: =3D and , > Both act as they do in the C language syntax. > > '=3D' is an assignment operator, taking a variable name on the left and > an expression on the right. > > ',' is a serial evaluation operator separating independently evaluatied > expressions. > > Combining these operators allows functions or expressions to have > side effects. In particular, a function or expression evaluation can > have the side effect of storing its result in a named variable. > > An example of using these to track successive data points is provided > in the new demo file data_feedback.dem > http://gnuplot.sourceforge.net/demo_4.3/data_feedback.html > Thanks, Ethan! Now I suppose I have to read a concise book on "C" to familiarise myself = with this syntax :) Levente |
|
From: Ethan M. <merritt@u.washington.edu> - 2007-10-23 20:13:37
|
On Tuesday 23 October 2007 13:01, ln...@dr... wrote: > > My biggest problem for the moment is the following: I did not find anything > useful about 'assignment' in the help file (I have the most recent CVS version > of gnuplot, of course). How should I use it then? Is there a configure-time > switch needed to compile it into gnuplot? No, but it is a very recent feature. Peter <pl...@pi...> referred to it as an "assignment feature" because the original trial implementation was in the form of a separate function named assign(). In the end, the same functionality was incorporated by adding support for two standard operators: = and , Both act as they do in the C language syntax. '=' is an assignment operator, taking a variable name on the left and an expression on the right. ',' is a serial evaluation operator separating independently evaluatied expressions. Combining these operators allows functions or expressions to have side effects. In particular, a function or expression evaluation can have the side effect of storing its result in a named variable. An example of using these to track successive data points is provided in the new demo file data_feedback.dem http://gnuplot.sourceforge.net/demo_4.3/data_feedback.html -- Ethan A Merritt Courier Deliveries: 1959 NE Pacific Dept of Biochemistry Health Sciences Building University of Washington - Seattle WA 98195-7742 |
|
From: <ln...@dr...> - 2007-10-23 20:01:44
|
> On Tue, 23 Oct 2007 18:05:29 +0200, <ln...@dr...> wrote: > >> Concerning the spline differentiating question: I am aware that spline= s >> can cause unwanted oscillations, > > Yes one of the problems with splines is that the contraints placed on > continuity and the fact that it must pass through all data points can l= ead > to some fairly unexpected and sometimes extreme excursions from the zon= e > where the real data lie. > I beg to disagree: you certainly are not forced to go through all your da= ta points for a smoothed spline. Of course, csplines do this way, but acspli= nes not (and bezier curves also do not necessarily pass through all the points). > Sometimes it's just clearly the wrong way to respresent certain data. > Sometimes it's less obvious but still wrong, this is more dangerous. Si= nce > the constraints on the continuity of the differencial are *completely > artificially* imposed as being the very definition of the spline fit , = I > believe it is utterly wrong to place any scientific meaning on the > derivative of a spline fit. > Did you try acsplines (it seems me what you say applies mostly for csplin= es)? I wouldn't say that the differential of such an acspline (if it is carefull= y drawn) has no scientific meaning. After all, you certainly remember the t= ime there were no personal computers and everybody was forced to draw the cur= ve himself for a dataset on a piece of paper then graphically differentiate = this curve by placing tangents and calculating their slopes. This is exactly t= he same thing I would like to do, but now aided by the computer. (Maybe smoothing= is not the exact term for this, that's why we seem to disagree?) Every measureme= nt, fitting, analysing method has its own inaccuracies and problems in the re= al world you must be aware of. This applies either to curve fitting or smoot= hing. It is only in mathematics where things are clearly defined and rigorous. > I would invite you to reflect on what that means for your results and > maybe have a quick look at the maths behind spline fitting to check wha= t I > say. The maths is pretty simple actually but it's important to know wha= t > you are doing in applying such a fit. > > Many people apply even least squares fits to completely inappropriate d= ata > because they do not realise the basic assumptions of the maths behind i= t, > that is that y errors are >> x errors. If that's not true you get the > wrong fit! > >> but in the past I was succesfully using >> acsplines with weights carefully chosen to approximate the curve I wou= ld >> draw by hand through the experimental data. > > Well I certainly hope you're not using that sort of bad science for > anything more than personal curiosity. > It does not sound like valid , reproducable scientific analysis. :) > I used this method to calculate meaningful results. At least not less mea= ningful that I would have obtained by graphical differentiation on a piece of pap= er. See: you have a bunch of data points and you simply don't know which is t= he mathematical function you should fit them to. You sometimes even don't ex= actly know what the exact error on x or y values is, but have only a clue. Don'= t misunderstand me: I don't need to have something very precise, only at gi= ven extent (certainly more precise than doing a "differentiation" on the poin= ts themselves, which have a certain scatter due to imprecisions and perturba= tions). By the way, even if I fit a mathematically well-defined function to these= data with the Marquardt-Levenberg method, I am not sure that this is the funct= ion which describes most accurately the phenomenon solely because the curve h= as a similar shape than the points plotted. > gnuplot claims not to do data processing (other than fitting user defin= ed > functions) although "smoothing" data with splines clearly is just that. > > I think you should be looking at applying some recognised data processi= ng > techniques to your data rather than hand tweeking splines. > > I hope my comments have highlighted some of the traps , in particular w= ith > regards to the derivative that you are interested in. > > regards, Peter. > My biggest problem for the moment is the following: I did not find anythi= ng useful about 'assignment' in the help file (I have the most recent CVS ve= rsion of gnuplot, of course). How should I use it then? Is there a configure-ti= me switch needed to compile it into gnuplot? Levente |
|
From: <pl...@pi...> - 2007-10-23 19:36:21
|
On Tue, 23 Oct 2007 17:09:10 +0200, Philipp K. Janert <ja...@ie...> wrote: > I can't really achieve this effect through a power law. > The gaussian that I suggested, however, allows this. > But a Lorentzian (1/(1+(d/a)**2)) would also do. As you said the gaussian would be an obvious choice , having a finite weight close to zero seems essencial. > Did you take a look at the images I posted? > (www.philipp-janert.com/gnuplot) I think the > Gaussian results in much more natural > approximations than the current > implementation does. Nice plots , I think you illustrate the point very well . There are some obvious artifacts in the current model particularly heavy in the lower order plots. I think your suggestion is a marked improvement from a theoretical point of view and that your example points out that the magnitude of the improvements possible by applying a more suitable weighting function. Nice contribution. I hope it gets adopted. /Peter. |
|
From: <HBB...@t-...> - 2007-10-23 19:19:26
|
Philipp K. Janert wrote: > I still don't think we are really on the same > page about this. > > We agree that dgrid3d calculates a weighted > average: > z = Sum_i w_i(d) z_i / Sum_i w_i(d) > where > z_i is the i-th raw data point > w_i(d) is the weight function Let's make that w(d_i). It's the distance d that depends on i, not the function that turns a distance into a w. We at least partly disagree at this point. The actual weight function is not w(d_i), but W_i := w(d_i) / Sum_j w(d_j) This is smooth (continuous even at d_i = 0), and it's limited to [0:1] as weights should be. > The current implementation uses > w(d) = 1/d**p if d > 0 > w(d) = 1 if d == 0 No, it doesn't. It's not really w(d), but the final W_i that's set to 1. This is achieved indirectly, by breaking out of the summation loop, such that the result becomes the first input point with d==0. > The lesser reason is this: > For d small, we multiply values by a large coefficient, > form the sum, then divide the large coefficient > out again. Numerically this is not very nice, because > we loose significant digits if we multiply with something, > sum, and then divide it out again. Multiplication doesn't lose digits at all, and summation only does where it has to, irrespective of scale. The multiply/divide pair has effectively no effect on numerical accuracy at all. In other words, the given algorithm is quite unlikely to have any bigger problem with round-off than a gaussian filter kernel would. |
|
From: <pl...@pi...> - 2007-10-23 19:11:44
|
On Tue, 23 Oct 2007 18:05:29 +0200, <ln...@dr...> wrote: > Concerning the spline differentiating question: I am aware that splines > can cause unwanted oscillations, Yes one of the problems with splines is that the contraints placed on continuity and the fact that it must pass through all data points can lead to some fairly unexpected and sometimes extreme excursions from the zone where the real data lie. Sometimes it's just clearly the wrong way to respresent certain data. Sometimes it's less obvious but still wrong, this is more dangerous. Since the constraints on the continuity of the differencial are *completely artificially* imposed as being the very definition of the spline fit , I believe it is utterly wrong to place any scientific meaning on the derivative of a spline fit. I would invite you to reflect on what that means for your results and maybe have a quick look at the maths behind spline fitting to check what I say. The maths is pretty simple actually but it's important to know what you are doing in applying such a fit. Many people apply even least squares fits to completely inappropriate data because they do not realise the basic assumptions of the maths behind it, that is that y errors are >> x errors. If that's not true you get the wrong fit! > but in the past I was succesfully using > acsplines with weights carefully chosen to approximate the curve I would > draw by hand through the experimental data. Well I certainly hope you're not using that sort of bad science for anything more than personal curiosity. It does not sound like valid , reproducable scientific analysis. :) gnuplot claims not to do data processing (other than fitting user defined functions) although "smoothing" data with splines clearly is just that. I think you should be looking at applying some recognised data processing techniques to your data rather than hand tweeking splines. I hope my comments have highlighted some of the traps , in particular with regards to the derivative that you are interested in. regards, Peter. |
|
From: <ln...@dr...> - 2007-10-23 16:05:35
|
On Mon, 2007-10-22 at 13:51 +0200, Juergen Wieferink wrote: > Well, AFAICS it has become possible in 4.3, though not very easily > accessible. > > There is a new assignment operator (and a comma operator), so you > can plot something like: > > old_x =3D NaN > plot 'test.dat' using 1:(dx=3D$1-old_x,old_x=3D$1,dx) > > This only works with data file plotting. You can use the '+' pseudo > file for fitted data (though I would recommend the analytic > derivative in this case). For smoothed curves, plot the results > into a table file (help set table) and use this file as given > above. > > -- Completely untested -- > > Juergen On Mon, 2007-10-22 at 16:26 +0200, pl...@pi... wrote: > Delta y / Delta x =3D y(row+1)-y(row)/(x(row+1)-x(row)). > > sorry I did reply to this one but it went to Levente only not the list. > > this sort of trivial task lends itself nicely to the new assign feature > (only in CVS AFAIK) and is exactly the sort of thing I had in mind when= I > requested it. > > check out the syntax since it is still being developed and may change. > Thanks J=FCrgen and Peter! I will try the assignment operator, it sounds promising. This was exactly my impression, that something new was introduced to gnuplot, allowing for (sort of) simple differentiation, only I did not remember exactly which command. Concerning the spline differentiating question: I am aware that splines can cause unwanted oscillations, but in the past I was succesfully using acsplines with weights carefully chosen to approximate the curve I would draw by hand through the experimental data. Sometimes it is faster this way than trying to find a good function and fit it with the Marquardt-Levenberg method (I often don't know exactly which function describes the experimental behaviour -- e.g. in complex biological systems). Levente |
|
From: Philipp K. J. <ja...@ie...> - 2007-10-23 15:09:16
|
I still don't think we are really on the same page about this. We agree that dgrid3d calculates a weighted average: z = Sum_i w_i(d) z_i / Sum_i w_i(d) where z_i is the i-th raw data point w_i(d) is the weight function w_i(d) is a function of d, which is the distance between the i-th raw data point and the current grid point. The current implementation uses w(d) = 1/d**p if d > 0 w(d) = 1 if d == 0 That means that w(d), as a function of d, is not continuous at d = 0. (I am not saying that the resulting approximation is not smooth as a function of x and y, as you point out. That is a different question, however.) The bigger question is whether this matters all that much. There are two reasons I don't find the current implementation optimal. The lesser reason is this: For d small, we multiply values by a large coefficient, form the sum, then divide the large coefficient out again. Numerically this is not very nice, because we loose significant digits if we multiply with something, sum, and then divide it out again. But since dgrid3d tries to calculate a rough-but-useful approximation, this is not so big of a problem. The bigger reason is that I would like to be able to control the range over which data points contribute to a grid point average! Ideally, dgrid3d would have a parameter, which controls the averaging range smoothly, such that as parameter -> 0 dgrid3d yields the raw data parameter -> infity dgrid3d averages over all points I can't really achieve this effect through a power law. The gaussian that I suggested, however, allows this. But a Lorentzian (1/(1+(d/a)**2)) would also do. Did you take a look at the images I posted? (www.philipp-janert.com/gnuplot) I think the Gaussian results in much more natural approximations than the current implementation does. Best, Ph. On Monday 22 October 2007 14:44, you wrote: > Philipp K. Janert wrote: > > But why a weight function w_i that is > > 1) discontinuous > > It's not really discontinuous. It's just undefined at d=0, but has a > continuous continuation (by w=1). This is really just a coordinate > singularity. > > > 2) grows above all bounds? > > Because it doesn't. The numerator and denominator are unlimited each by > itself, but the quotient, which is the real weight factor, is limited to > the interval [0:1], just like it should be. > > > If nothing else, numerically this is not > > good (round-off). But it also does not > > seem to be right that if a grid point > > coincides exactly with a data point, it > > gets weight 1, and weight 1/eps >> 1 > > if the grid point is shifted by a very small > > amount eps << 1 from the grid point. > > Which is why that's _not_ what we actually do. The normalization takes > care of that. The real weight of a point very close to the sample is > very close to 1, because its weight will be much larger than all the > others, so the sum(w_i) will be dominated by it, and those two large > terms cancel each other. > > >> I don't think there should be any such radius at all, in the typical > >> case. The algorithm should be independent of the x and y range of its > >> input, if at all possible. > > > > Exactly. But the current one isn't. If x values > > are spaced by 100, and y values spaced by 1 > > in x and y units, the weight function will behave > > differently in both directions. > > Not really. It will _look_ like it behaved differently, because the > axis ranges are so different. But that's a plotting artifact, not a > property of dgrid3d. The algorithm operates on actual data values > rather than plot coordinates, and it treats those exactly equally. > > If and when that's a problem, data can always be re-scaled before being > presented to dgrid3d. |
|
From: Ethan M. <merritt@u.washington.edu> - 2007-10-22 23:19:25
|
On Monday 22 October 2007 07:06, Fr=E9d=E9ric Gava wrote: > set autoscale > set terminal epslatex color dashed "default" 13 > set output "test.eps" > set xlabel "Number n" > set ylabel "(s)" > set title "Computation" > set key left > plot 'test.data' using 1:2 title "Test" with linespoints It works fine for me. I attach a gzipped test.dvi file produced by running the output of your test commands through LaTeX. =2D-=20 Ethan A Merritt |
|
From: Ethan M. <merritt@u.washington.edu> - 2007-10-22 17:27:16
|
The attached file did not make it through to the mailing list, so I can not try it out here. Please open a bug report on the SourceForge site and upload the problematic data file there for testing. http://sourceforge.net/tracker/?group_id=2055&atid=102055 On Monday 22 October 2007 04:40, Lars Hecking wrote: > ----- Forwarded message from Jonathan Lawson <jon...@bi...> ----- > > From: Jonathan Lawson > Sent: Thursday, October 11, 2007 11:39 AM > To: 'lhe...@nm...' > Subject: Gnuplot .. > > Hi Lars, > > Trust all remains well with you ... I was trying to plot data > from a CSV file, which I did with the following difficulty: > > If there was a blank space following the commas in the file, there was no > problem; however if there are no white spaces following the commas then > gnuplot reports that it contains no valid data. But it does this only on > complex files: A simple file with comma separated data seems to work fine. I > have attached the file in question and the command which I am using to plot > the data is > > > > Plot "Filename" u 3 > > > > Where Filename is either one of the files attached to this email. -- Ethan A Merritt |
|
From: <pl...@pi...> - 2007-10-22 14:26:23
|
Delta y / Delta x =3D y(row+1)-y(row)/(x(row+1)-x(row)).
sorry I did reply to this one but it went to Levente only not the list.
this sort of trivial task lends itself nicely to the new assign feature =
=
(only in CVS AFAIK) and is exactly the sort of thing I had in mind when =
I =
requested it.
check out the syntax since it is still being developed and may change.
It does not enable you to look ahead or backwards in the data but settin=
g =
a variables like last_x last_y and diffing to the previous point rather =
=
than the next will do this nicely.
This can be done inline in the plot command or via a function call in th=
e =
plot command , that keeps the plot line less cluttered.
Here's two examples: sample max x and y values in plot range and =
trapezoidal integration area under graph. Again check syntax since I thi=
nk =
it has changed since I built this CVS snapshot.
xmax=3D0;ymax=3D0;
max(x,y)=3Dassign("ymax",(ymax<y)?y+0. * assign("xmax",x):ymax );
started=3D0;
add_aug(x,y)=3Dassign("area",(started>0)?\
area + (x-prev_x)*(y+prev_y)/2.0 =
=
+ 0.0*(assign("prev_x",x) \
+assign("prev_y",y) ) \
: 0.0*( assign("started",1) =
+ assign("prev_x",x) + assign("prev_y",y) ) \
) ;
this using clause is in my plot command:
using 1:(add_aug( ($1),($2) ))
HTH, Peter.
|
|
From: <ga...@un...> - 2007-10-22 14:06:11
|
Dear "Gnuplot developer" I give the following commands to Gnuplot : set autoscale set terminal epslatex color dashed "default" 13 set output "test.eps" set xlabel "Number n" set ylabel "(s)" set title "Computation" set key left plot 'test.data' using 1:2 title "Test" with linespoints but I have a big bug: the keys are not in the left but in the center (no problem with "set key right"). Best regards, Frederic Gava |
|
From: Juergen W. <wie...@fr...> - 2007-10-22 11:51:45
|
> AFAIK calculating or approximating the 1st order derivative of a > function or x-y dataset was not (easily?) possible a few years ago from > within gnuplot, but maybe this has changed since then -- I don't follow > closely gnuplot's development anymore. > > My problem is the following: from a dataset, I first plot the x-y > values, then either fit a function to them or make a smoothing with > acsplines. So far, no problem. But afterwards, I need the 1st order > derivative of the fitted or smoothed curve. This would be easy if > gnuplot had a way to access data not only from the _actual_ row but also > from the following one in order to calculate Delta y / Delta x = y(row > +1)-y(row)/(x(row+1)-x(row)). This was not possible earlier. Has the > situation changed since then? If not, is it still possible to perform > differentiation from within gnuplot (even if this means using some awk > script or similar), as I would like to automate the process and gnuplot > is easily driven from a shell script or batch file. Well, AFAICS it has become possible in 4.3, though not very easily accessible. There is a new assignment operator (and a comma operator), so you can plot something like: old_x = NaN plot 'test.dat' using 1:(dx=$1-old_x,old_x=$1,dx) This only works with data file plotting. You can use the '+' pseudo file for fitted data (though I would recommend the analytic derivative in this case). For smoothed curves, plot the results into a table file (help set table) and use this file as given above. -- Completely untested -- Juergen |