You can subscribe to this list here.
| 2001 |
Jan
|
Feb
(1) |
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
|
Sep
|
Oct
|
Nov
|
Dec
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2002 |
Jan
(1) |
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
|
Nov
(1) |
Dec
|
| 2003 |
Jan
|
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
(83) |
Nov
(57) |
Dec
(111) |
| 2004 |
Jan
(38) |
Feb
(121) |
Mar
(107) |
Apr
(241) |
May
(102) |
Jun
(190) |
Jul
(239) |
Aug
(158) |
Sep
(184) |
Oct
(193) |
Nov
(47) |
Dec
(68) |
| 2005 |
Jan
(190) |
Feb
(105) |
Mar
(99) |
Apr
(65) |
May
(92) |
Jun
(250) |
Jul
(197) |
Aug
(128) |
Sep
(101) |
Oct
(183) |
Nov
(186) |
Dec
(42) |
| 2006 |
Jan
(102) |
Feb
(122) |
Mar
(154) |
Apr
(196) |
May
(181) |
Jun
(281) |
Jul
(310) |
Aug
(198) |
Sep
(145) |
Oct
(188) |
Nov
(134) |
Dec
(90) |
| 2007 |
Jan
(134) |
Feb
(181) |
Mar
(157) |
Apr
(57) |
May
(81) |
Jun
(204) |
Jul
(60) |
Aug
(37) |
Sep
(17) |
Oct
(90) |
Nov
(122) |
Dec
(72) |
| 2008 |
Jan
(130) |
Feb
(108) |
Mar
(160) |
Apr
(38) |
May
(83) |
Jun
(42) |
Jul
(75) |
Aug
(16) |
Sep
(71) |
Oct
(57) |
Nov
(59) |
Dec
(152) |
| 2009 |
Jan
(73) |
Feb
(213) |
Mar
(67) |
Apr
(40) |
May
(46) |
Jun
(82) |
Jul
(73) |
Aug
(57) |
Sep
(108) |
Oct
(36) |
Nov
(153) |
Dec
(77) |
| 2010 |
Jan
(42) |
Feb
(171) |
Mar
(150) |
Apr
(6) |
May
(22) |
Jun
(34) |
Jul
(31) |
Aug
(38) |
Sep
(32) |
Oct
(59) |
Nov
(13) |
Dec
(62) |
| 2011 |
Jan
(114) |
Feb
(139) |
Mar
(126) |
Apr
(51) |
May
(53) |
Jun
(29) |
Jul
(41) |
Aug
(29) |
Sep
(35) |
Oct
(87) |
Nov
(42) |
Dec
(20) |
| 2012 |
Jan
(111) |
Feb
(66) |
Mar
(35) |
Apr
(59) |
May
(71) |
Jun
(32) |
Jul
(11) |
Aug
(48) |
Sep
(60) |
Oct
(87) |
Nov
(16) |
Dec
(38) |
| 2013 |
Jan
(5) |
Feb
(19) |
Mar
(41) |
Apr
(47) |
May
(14) |
Jun
(32) |
Jul
(18) |
Aug
(68) |
Sep
(9) |
Oct
(42) |
Nov
(12) |
Dec
(10) |
| 2014 |
Jan
(14) |
Feb
(139) |
Mar
(137) |
Apr
(66) |
May
(72) |
Jun
(142) |
Jul
(70) |
Aug
(31) |
Sep
(39) |
Oct
(98) |
Nov
(133) |
Dec
(44) |
| 2015 |
Jan
(70) |
Feb
(27) |
Mar
(36) |
Apr
(11) |
May
(15) |
Jun
(70) |
Jul
(30) |
Aug
(63) |
Sep
(18) |
Oct
(15) |
Nov
(42) |
Dec
(29) |
| 2016 |
Jan
(37) |
Feb
(48) |
Mar
(59) |
Apr
(28) |
May
(30) |
Jun
(43) |
Jul
(47) |
Aug
(14) |
Sep
(21) |
Oct
(26) |
Nov
(10) |
Dec
(2) |
| 2017 |
Jan
(26) |
Feb
(27) |
Mar
(44) |
Apr
(11) |
May
(32) |
Jun
(28) |
Jul
(75) |
Aug
(45) |
Sep
(35) |
Oct
(285) |
Nov
(99) |
Dec
(16) |
| 2018 |
Jan
(8) |
Feb
(8) |
Mar
(42) |
Apr
(35) |
May
(23) |
Jun
(12) |
Jul
(16) |
Aug
(11) |
Sep
(8) |
Oct
(16) |
Nov
(5) |
Dec
(8) |
| 2019 |
Jan
(9) |
Feb
(28) |
Mar
(4) |
Apr
(10) |
May
(7) |
Jun
(4) |
Jul
(4) |
Aug
|
Sep
(4) |
Oct
|
Nov
(23) |
Dec
(3) |
| 2020 |
Jan
(19) |
Feb
(3) |
Mar
(22) |
Apr
(17) |
May
(10) |
Jun
(69) |
Jul
(18) |
Aug
(23) |
Sep
(25) |
Oct
(11) |
Nov
(20) |
Dec
(9) |
| 2021 |
Jan
(1) |
Feb
(7) |
Mar
(9) |
Apr
|
May
(1) |
Jun
(8) |
Jul
(6) |
Aug
(8) |
Sep
(7) |
Oct
|
Nov
(2) |
Dec
(23) |
| 2022 |
Jan
(23) |
Feb
(9) |
Mar
(9) |
Apr
|
May
(8) |
Jun
(1) |
Jul
(6) |
Aug
(8) |
Sep
(30) |
Oct
(5) |
Nov
(4) |
Dec
(6) |
| 2023 |
Jan
(2) |
Feb
(5) |
Mar
(7) |
Apr
(3) |
May
(8) |
Jun
(45) |
Jul
(8) |
Aug
|
Sep
(2) |
Oct
(14) |
Nov
(7) |
Dec
(2) |
| 2024 |
Jan
(4) |
Feb
(4) |
Mar
|
Apr
(7) |
May
(2) |
Jun
(1) |
Jul
|
Aug
(5) |
Sep
|
Oct
|
Nov
(4) |
Dec
(14) |
| 2025 |
Jan
(22) |
Feb
(6) |
Mar
(5) |
Apr
(14) |
May
(6) |
Jun
(11) |
Jul
(19) |
Aug
|
Sep
(17) |
Oct
(1) |
Nov
(2) |
Dec
(18) |
| 2026 |
Jan
|
Feb
|
Mar
(5) |
Apr
|
May
(2) |
Jun
(1) |
Jul
(6) |
Aug
(1) |
Sep
|
Oct
|
Nov
|
Dec
|
|
From: sfeam (E. Merritt) <eam...@gm...> - 2011-05-24 03:35:36
|
Christoph Bersch <us...@be...> wrote> > > is there a specific reason why the 'transparent' terminal option applies > only to the pngcairo and not to the pdfcairo terminal? > > Christoph So far as I know, pdf and PostScript files have no "background" to be set transparent or opaque. When you print the file, the color of the paper it is printed on becomes the background color. Do you actually see a difference in the output file after your patch? Using what tool to view it? Ethan |
|
From: Ethan M. <merritt@u.washington.edu> - 2011-05-24 00:08:17
|
On Monday, May 23, 2011 01:12:23 am pl...@pi... wrote: > > if there is an alternative approach for the case when x > > uncertainty can't be ignored, keep the discussion going and maybe we can > > add such a feature to the list of items we'd like to add in the future. > > > > Dan > > Yes , I would love to propose that. Some kind of total least squares may > be an option I don't know enough about that to make a concrete proposal. It is standard in my field to use instead a maximum likelihood residual that allows for separate error distributions on x and y. The maximum likelihood treatment reduces to a weighted least squares treatment if and only if the distribution of errors on both x and y are (1) Gaussian and (2) of equal magnitude. Gnuplot allows for input of precalculated non-uniform weights on y, which partially addresses (2). But there is no provision for non-Gaussian errors on y, and no provision for any error model at all on x. If the errors in your data do not follow the simple Gaussian model, then almost certainly you would be better off using maximum likelihood rather than least-squares. But in order to do so you need first a model for the error distributions. That may be obvious for any particular experiment or source of data, but it's difficult to impossible for a general-purpose program to figure it out for you. You're the one who knows the source of the data, so it's up to you to provide an appropriate explicit error model along with the data. The point is that switching to a better minimization residual is more than just a matter of changing the internal code. It would require a more complex description of your data that includes error models for both the independent and dependent variables. NB: When I say "x", I really mean the full set of independent variables [x1,x2,x3,...]. "y" is the single dependent variable estimated by f(x1,x2,x3,...) Ethan -- Ethan A Merritt Biomolecular Structure Center, K-428 Health Sciences Bldg University of Washington, Seattle 98195-7742 |
|
From: Christoph B. <us...@be...> - 2011-05-23 12:19:08
|
Hi, is there a specific reason why the 'transparent' terminal option applies only to the pngcairo and not to the pdfcairo terminal? I removed the two respective lines in term/cairo.trm and it worked well for me: --- gnuplot-cvs/term/cairo.trm 2011-05-23 13:48:46.000000000 +0200 +++ gnuplot/term/cairo.trm 2011-05-23 14:00:16.000000000 +0200 @@ -299,13 +299,11 @@ break; case CAIROTRM_TRANSPARENT: c_token++; - if (!strcmp(term->name,"pngcairo")) - cairo_params->transparent = TRUE; + cairo_params->transparent = TRUE; break; case CAIROTRM_NOTRANSPARENT: c_token++; - if (!strcmp(term->name,"pngcairo")) - cairo_params->transparent = FALSE; + cairo_params->transparent = FALSE; break; case CAIROTRM_CROP: c_token++; I only tested it with set terminal pdfcairo transparent set output 'test-transparency.pdf' plot sin(x) But because there is an explicit test for the pngcairo terminal, I suppose there was a good reason for this? Christoph |
|
From: <pl...@pi...> - 2011-05-23 08:39:02
|
On 05/23/11 02:55, Daniel J Sebald wrote: > On 05/22/2011 01:14 PM, pl...@pi... wrote: >> On 05/21/11 23:04, Daniel J Sebald wrote: >>> On 05/21/2011 03:15 AM, pl...@pi... wrote: >>>> Hi, >>>> >>>> 'help fit' reports that the fit command uses Levenberg–Marquardt algo to >>>> do the fit. >>>> >>>> I think this raises an important question that very often over-looked by >>>> many users of least-squares techniques even at maths PhD level. >>>> >>>> Such techniques often only optimised the y error rather than the >>>> perpendicular error from the line. This is implicitly assuming y >>>> uncertainty>> x uncertainty. While this condition is often satisfied >>>> in a controlled experiment there are many situations where this is not >>>> applicable and gets totally overlooked. >>>> >>>> A common case is scatter plots which are frequently used to seek a >>>> relations between two quantities , each with significant errors / >>>> uncertainties. >>>> >>>> In this situation the fitted line is "wrong". In fact it's the >>>> application that is wrong , hence the wrong result. This may or may not >>>> be apparent to the eye. >>>> >>>> I have seen this happen so many times (including once in a PhD thesis >>>> report!) that I think it needs a serious health warning in the doc. >>>> >>>> "Warning: using least-squares inappropriately can seriously damage your >>>> reputation". ;) >>>> >>>> Firstly , could you confirm the basis on which this algo is applied in >>>> gnuplot? Does it only optimise vertical y residuals? >>> >>> Often it is an assumption that the independent variables are exact >>> measurements. Not true, typically, but if the variance is small and >>> homoscedastic, the two can probably be lumped together. I.e., we are >>> searching for a relationship: >>> >>> Y = f(X + eps1) + eps2 >>> ~= f(X) + C eps1 + residual + eps2 >>> ~= f(X) + (C eps1 + eps2) >>> >>> where hopefully the residual due to nonlinearity of the relationship is >>> small compared to other randomness. It's up to the user's judgment and >>> knowledge of the application to determine that. >>> >>> Anyway, your point is true of most software packages: details are so >>> often lacking. That's why it would be nice to have a set of white >>> papers to go along with the software so that people know exactly what >>> the algorithm is, both for the benefit of the user and other developers. >>> Most of the time it is "here's a hunk of code, use it at your own risk". >>> >>> Dan >>> >> >> Thanks, I've read up on that algo and clearly this is only doing NLLS on >> y residuals. >> >> So my suggestion is that this is made abundantly clear in the help text. >> I'm not suggesting your "while paper" but just some comment to the >> effect that 'fit' will not give correct results if there are non >> negligible errors in x values. >> >> It absolutely amazes me how few people realise this , even highly >> qualified ones, so this is not some pedantic nicety. >> >> Most people seem to think once they've heard of doing a least squares >> fit that's all there is to it and it's some magical formula that works >> for all cases. >> >> I suggest modifying the first paragraph of help fit with something like >> the following: >> >> >> >> The `fit` command can fit a user-supplied expression to a set of data >> points >> (x,z) or (x,y,z), using an implementation of the nonlinear least-squares >> (NLLS) Marquardt-Levenberg algorithm. Any user-defined variable >> occurring in >> the expression may serve as a fit parameter, but the return type of the >> expression must be real. >> >> >> >> new >> >> >> The `fit` command can fit a user-supplied expression to a set of data >> points >> (x,z) or (x,y,z), using an implementation of the non-linear least-squares >> (NLLS) Marquardt-Levenberg algorithm. This algorithm optimises y >> residuals only and >> carries the implicit assumption that error/uncertainties in x are >> negligible. If that is not the case the fit may succeed but will give >> wrong results. > > I'm OK with the change, but for a few things. > > First, you stated 'optimizes y' and 'uncertainties in x', but please > check that this precisely describes the algorithm. The beginning of the > paragraph lists (x, z) or (x, y, z). The first expression has no 'y', > so what is optimized in that case? The second expression has (x, y, z), > so is it only the 'x' in that case that is assumed exact? Or is it both > 'x' and 'y'? Perhaps that sentence should be, "This algorithm optimises > z residuals only and carries the implicit assumption that > error/uncertainties in x (and y) are negligible." very good point , I was trying to keep it brief (and perhaps dumb it down) but it should correctly refer to dependent and independent variables. I was concerned that may mean the warning was lost on many users. Perhaps "independant variable (eg. x axis)" or similar would be better. > > Second, I would hold off using the statement "wrong results" and maybe > use "a poor fit". Saying the result is wrong means the algorithm is > broken, but it just provides numbers. The user is using the wrong tool. > Also, "wrong" means a bad fit in this context and one can get a bad > fit and misinterpret results even if uncertainty in x is small. No, I think wrong result is exactly the case. Wrong result does not mean the algo is "broken" it means the wrong tool was used and the result obtained was wrong. The result is not "poor" it is wrong. Watering down the language does not shift the blame. Maybe some wording like "wrong technique" could help underline the cause of the problem but I think "wrong" definitely needs to be in there. There's no hedging around the fact, it can be considerably wrong. > > Third, if there is an alternative approach for the case when x > uncertainty can't be ignored, keep the discussion going and maybe we can > add such a feature to the list of items we'd like to add in the future. > > Dan > Yes , I would love to propose that. Some kind of total least squares may be an option I don't know enough about that to make a concrete proposal. I suspect it may be too variable to include as a turn key option like fit but I would really like that option if it is possible. thanks for your thoughtful comments. |
|
From: Daniel J S. <dan...@ie...> - 2011-05-23 00:55:19
|
On 05/22/2011 01:14 PM, pl...@pi... wrote: > On 05/21/11 23:04, Daniel J Sebald wrote: >> On 05/21/2011 03:15 AM, pl...@pi... wrote: >>> Hi, >>> >>> 'help fit' reports that the fit command uses Levenberg–Marquardt algo to >>> do the fit. >>> >>> I think this raises an important question that very often over-looked by >>> many users of least-squares techniques even at maths PhD level. >>> >>> Such techniques often only optimised the y error rather than the >>> perpendicular error from the line. This is implicitly assuming y >>> uncertainty>> x uncertainty. While this condition is often satisfied >>> in a controlled experiment there are many situations where this is not >>> applicable and gets totally overlooked. >>> >>> A common case is scatter plots which are frequently used to seek a >>> relations between two quantities , each with significant errors / >>> uncertainties. >>> >>> In this situation the fitted line is "wrong". In fact it's the >>> application that is wrong , hence the wrong result. This may or may not >>> be apparent to the eye. >>> >>> I have seen this happen so many times (including once in a PhD thesis >>> report!) that I think it needs a serious health warning in the doc. >>> >>> "Warning: using least-squares inappropriately can seriously damage your >>> reputation". ;) >>> >>> Firstly , could you confirm the basis on which this algo is applied in >>> gnuplot? Does it only optimise vertical y residuals? >> >> Often it is an assumption that the independent variables are exact >> measurements. Not true, typically, but if the variance is small and >> homoscedastic, the two can probably be lumped together. I.e., we are >> searching for a relationship: >> >> Y = f(X + eps1) + eps2 >> ~= f(X) + C eps1 + residual + eps2 >> ~= f(X) + (C eps1 + eps2) >> >> where hopefully the residual due to nonlinearity of the relationship is >> small compared to other randomness. It's up to the user's judgment and >> knowledge of the application to determine that. >> >> Anyway, your point is true of most software packages: details are so >> often lacking. That's why it would be nice to have a set of white >> papers to go along with the software so that people know exactly what >> the algorithm is, both for the benefit of the user and other developers. >> Most of the time it is "here's a hunk of code, use it at your own risk". >> >> Dan >> > > Thanks, I've read up on that algo and clearly this is only doing NLLS on > y residuals. > > So my suggestion is that this is made abundantly clear in the help text. > I'm not suggesting your "while paper" but just some comment to the > effect that 'fit' will not give correct results if there are non > negligible errors in x values. > > It absolutely amazes me how few people realise this , even highly > qualified ones, so this is not some pedantic nicety. > > Most people seem to think once they've heard of doing a least squares > fit that's all there is to it and it's some magical formula that works > for all cases. > > I suggest modifying the first paragraph of help fit with something like > the following: > > >> > The `fit` command can fit a user-supplied expression to a set of data > points > (x,z) or (x,y,z), using an implementation of the nonlinear least-squares > (NLLS) Marquardt-Levenberg algorithm. Any user-defined variable > occurring in > the expression may serve as a fit parameter, but the return type of the > expression must be real. > >> > > new > >> > The `fit` command can fit a user-supplied expression to a set of data > points > (x,z) or (x,y,z), using an implementation of the non-linear least-squares > (NLLS) Marquardt-Levenberg algorithm. This algorithm optimises y > residuals only and > carries the implicit assumption that error/uncertainties in x are > negligible. If that is not the case the fit may succeed but will give > wrong results. I'm OK with the change, but for a few things. First, you stated 'optimizes y' and 'uncertainties in x', but please check that this precisely describes the algorithm. The beginning of the paragraph lists (x, z) or (x, y, z). The first expression has no 'y', so what is optimized in that case? The second expression has (x, y, z), so is it only the 'x' in that case that is assumed exact? Or is it both 'x' and 'y'? Perhaps that sentence should be, "This algorithm optimises z residuals only and carries the implicit assumption that error/uncertainties in x (and y) are negligible." Second, I would hold off using the statement "wrong results" and maybe use "a poor fit". Saying the result is wrong means the algorithm is broken, but it just provides numbers. The user is using the wrong tool. Also, "wrong" means a bad fit in this context and one can get a bad fit and misinterpret results even if uncertainty in x is small. Third, if there is an alternative approach for the case when x uncertainty can't be ignored, keep the discussion going and maybe we can add such a feature to the list of items we'd like to add in the future. Dan |
|
From: <pl...@pi...> - 2011-05-22 18:14:26
|
On 05/21/11 23:04, Daniel J Sebald wrote: > On 05/21/2011 03:15 AM, pl...@pi... wrote: >> Hi, >> >> 'help fit' reports that the fit command uses Levenberg–Marquardt algo to >> do the fit. >> >> I think this raises an important question that very often over-looked by >> many users of least-squares techniques even at maths PhD level. >> >> Such techniques often only optimised the y error rather than the >> perpendicular error from the line. This is implicitly assuming y >> uncertainty>> x uncertainty. While this condition is often satisfied >> in a controlled experiment there are many situations where this is not >> applicable and gets totally overlooked. >> >> A common case is scatter plots which are frequently used to seek a >> relations between two quantities , each with significant errors / >> uncertainties. >> >> In this situation the fitted line is "wrong". In fact it's the >> application that is wrong , hence the wrong result. This may or may not >> be apparent to the eye. >> >> I have seen this happen so many times (including once in a PhD thesis >> report!) that I think it needs a serious health warning in the doc. >> >> "Warning: using least-squares inappropriately can seriously damage your >> reputation". ;) >> >> Firstly , could you confirm the basis on which this algo is applied in >> gnuplot? Does it only optimise vertical y residuals? > > Often it is an assumption that the independent variables are exact > measurements. Not true, typically, but if the variance is small and > homoscedastic, the two can probably be lumped together. I.e., we are > searching for a relationship: > > Y = f(X + eps1) + eps2 > ~= f(X) + C eps1 + residual + eps2 > ~= f(X) + (C eps1 + eps2) > > where hopefully the residual due to nonlinearity of the relationship is > small compared to other randomness. It's up to the user's judgment and > knowledge of the application to determine that. > > Anyway, your point is true of most software packages: details are so > often lacking. That's why it would be nice to have a set of white > papers to go along with the software so that people know exactly what > the algorithm is, both for the benefit of the user and other developers. > Most of the time it is "here's a hunk of code, use it at your own risk". > > Dan > Thanks, I've read up on that algo and clearly this is only doing NLLS on y residuals. So my suggestion is that this is made abundantly clear in the help text. I'm not suggesting your "while paper" but just some comment to the effect that 'fit' will not give correct results if there are non negligible errors in x values. It absolutely amazes me how few people realise this , even highly qualified ones, so this is not some pedantic nicety. Most people seem to think once they've heard of doing a least squares fit that's all there is to it and it's some magical formula that works for all cases. I suggest modifying the first paragraph of help fit with something like the following: >> The `fit` command can fit a user-supplied expression to a set of data points (x,z) or (x,y,z), using an implementation of the nonlinear least-squares (NLLS) Marquardt-Levenberg algorithm. Any user-defined variable occurring in the expression may serve as a fit parameter, but the return type of the expression must be real. >> new >> The `fit` command can fit a user-supplied expression to a set of data points (x,z) or (x,y,z), using an implementation of the non-linear least-squares (NLLS) Marquardt-Levenberg algorithm. This algorithm optimises y residuals only and carries the implicit assumption that error/uncertainties in x are negligible. If that is not the case the fit may succeed but will give wrong results. Any user-defined variable occurring in the expression may serve as a fit parameter, but the return type of the expression must be real. >> regards, Peter. |
|
From: Daniel J S. <dan...@ie...> - 2011-05-21 21:04:51
|
On 05/21/2011 03:15 AM, pl...@pi... wrote: > Hi, > > 'help fit' reports that the fit command uses Levenberg–Marquardt algo to > do the fit. > > I think this raises an important question that very often over-looked by > many users of least-squares techniques even at maths PhD level. > > Such techniques often only optimised the y error rather than the > perpendicular error from the line. This is implicitly assuming y > uncertainty>> x uncertainty. While this condition is often satisfied > in a controlled experiment there are many situations where this is not > applicable and gets totally overlooked. > > A common case is scatter plots which are frequently used to seek a > relations between two quantities , each with significant errors / > uncertainties. > > In this situation the fitted line is "wrong". In fact it's the > application that is wrong , hence the wrong result. This may or may not > be apparent to the eye. > > I have seen this happen so many times (including once in a PhD thesis > report!) that I think it needs a serious health warning in the doc. > > "Warning: using least-squares inappropriately can seriously damage your > reputation". ;) > > Firstly , could you confirm the basis on which this algo is applied in > gnuplot? Does it only optimise vertical y residuals? Often it is an assumption that the independent variables are exact measurements. Not true, typically, but if the variance is small and homoscedastic, the two can probably be lumped together. I.e., we are searching for a relationship: Y = f(X + eps1) + eps2 ~= f(X) + C eps1 + residual + eps2 ~= f(X) + (C eps1 + eps2) where hopefully the residual due to nonlinearity of the relationship is small compared to other randomness. It's up to the user's judgment and knowledge of the application to determine that. Anyway, your point is true of most software packages: details are so often lacking. That's why it would be nice to have a set of white papers to go along with the software so that people know exactly what the algorithm is, both for the benefit of the user and other developers. Most of the time it is "here's a hunk of code, use it at your own risk". Dan |
|
From: <pl...@pi...> - 2011-05-21 08:15:37
|
Hi, 'help fit' reports that the fit command uses Levenberg–Marquardt algo to do the fit. I think this raises an important question that very often over-looked by many users of least-squares techniques even at maths PhD level. Such techniques often only optimised the y error rather than the perpendicular error from the line. This is implicitly assuming y uncertainty >> x uncertainty. While this condition is often satisfied in a controlled experiment there are many situations where this is not applicable and gets totally overlooked. A common case is scatter plots which are frequently used to seek a relations between two quantities , each with significant errors / uncertainties. In this situation the fitted line is "wrong". In fact it's the application that is wrong , hence the wrong result. This may or may not be apparent to the eye. I have seen this happen so many times (including once in a PhD thesis report!) that I think it needs a serious health warning in the doc. "Warning: using least-squares inappropriately can seriously damage your reputation". ;) Firstly , could you confirm the basis on which this algo is applied in gnuplot? Does it only optimise vertical y residuals? Thanks, Peter. |
|
From: Hans-Bernhard B. <HBB...@t-...> - 2011-05-18 22:03:26
|
On 14.05.2011 13:06, Petr Mikulik wrote: > A colleague has noticed wrong parameter values after 2D fitting via > fit f(x,y) 'data.txt' us 1:2:3 via a,b,c > > Gnuplot was fitting correctly after adding the dummy weight column as > "help fit" indicates: > fit f(x,y) 'data.txt' us 1:2:3:(1) via a,b,c > > I think gnuplot should write an error of not having enough columns or > missing weights column in the former command or it should fill weights by > one if the 4th column is missing. That idea is based on a misconception. The problem at hand is not that gnuplot didn't fill in the weights --- it's that the user incorrectly believed the first command given above were a 2-variable fit (somewhat incorrectly referred to as a "2D" one), which it's not. That's a perfectly valid command to perform a 1-variable fit (assuming a variable 'y' was defined before) that, and gnuplot has no way of knowing that's not what the user wanted. I.e. there is no "missing" column specification here, so nothing to diagnose. |
|
From: Petr M. <mi...@ph...> - 2011-05-14 11:06:17
|
A colleague has noticed wrong parameter values after 2D fitting via
fit f(x,y) 'data.txt' us 1:2:3 via a,b,c
Gnuplot was fitting correctly after adding the dummy weight column as
"help fit" indicates:
fit f(x,y) 'data.txt' us 1:2:3:(1) via a,b,c
I think gnuplot should write an error of not having enough columns or
missing weights column in the former command or it should fill weights by
one if the 4th column is missing.
---
PM
|
|
From: Bastian M. <bma...@we...> - 2011-05-13 14:56:58
|
Am 13.05.2011 14:35, schrieb pl...@pi...: >>> While the subject is being raised , may I suggest a couple of possible >>> feature improvements in this area? >>> >>> 1) This cycling of "mode" could also be done by clicking on the status >>> message (click is mapped to "1" key). This would avoid switching mouse >>> to keyboard which can be handy. >> >> Could be done, certainly, but there is currently no code to interpret the >> mouse location in terms like "on the status message". > > Well there is click on graph area (pause mouse etc) so I would have > thought x_bottom_graph and x_bottom_window would be close to what is > needed. Mouse events outside the graph area don't get passed on to the gnuplot kernel. It would be straight forward to translate those events to keyboard events, though. A trivial patch for Windows is attached. Bastian |
|
From: <pl...@pi...> - 2011-05-13 14:21:50
|
On 05/13/11 05:27, sfeam (Ethan Merritt) wrote: > On Thursday, 12 May 2011, pl...@pi... wrote: >> >> I had tried the h key help but had not appreciated what those variables >> referred to: >> >> 1 `builtin-decrement-mousemode` >> 2 `builtin-increment-mousemode` >> >> Maybe those terms could be more explicit, it's not clear what they refer >> to. >> >> The content of the various modes is equally a bit arcane and the help >> doesn't: >> >> >> This corresponds to the key >> >> bindings '1', '2', '3', '4' (see the driver's documentation) >> >> I find no detail for example in help wxt. What does this comment refer to? > > I was surprised to find that the documentation is seriously out of date, > as in "doesn't describe the actual behavior of any gnuplot version since > before the start of the CVS repository". I have now updated the 4.5 > documentation a bit. OK, thanks, glad you caught that one. > >> While the subject is being raised , may I suggest a couple of possible >> feature improvements in this area? >> >> 1) This cycling of "mode" could also be done by clicking on the status >> message (click is mapped to "1" key). This would avoid switching mouse >> to keyboard which can be handy. > > Could be done, certainly, but there is currently no code to interpret the > mouse location in terms like "on the status message". Well there is click on graph area (pause mouse etc) so I would have thought x_bottom_graph and x_bottom_window would be close to what is needed. > >> 2) As I understand it this only shows x1y1 coords. It would clearly be >> useful when using x2y2 as well if those were available in this read-out. >> One way to trigger this would be a click on the legend entry for a >> line or on/near the relevant axis. (some thought needed to avoid >> ambiguity with the new SVG toggle , pause mouse or other functions. ) > > Huh. For me the x2y2 axes are echoed properly if they are active in the > current plot. Tested on wxt, x11, and canvas. It is true that I ignored > x2y2 in the very recent svg mouse-tracking code. What terminal are you > using that fails to show them? Sorry , my incorrect recollection. It most likely was in relation to the SVG changes that I'd retained that idea. > > Ethan > |
|
From: sfeam (E. Merritt) <eam...@gm...> - 2011-05-13 03:27:45
|
On Thursday, 12 May 2011, pl...@pi... wrote: > > I had tried the h key help but had not appreciated what those variables > referred to: > > 1 `builtin-decrement-mousemode` > 2 `builtin-increment-mousemode` > > Maybe those terms could be more explicit, it's not clear what they refer > to. > > The content of the various modes is equally a bit arcane and the help > doesn't: > > >> This corresponds to the key > >> bindings '1', '2', '3', '4' (see the driver's documentation) > > I find no detail for example in help wxt. What does this comment refer to? I was surprised to find that the documentation is seriously out of date, as in "doesn't describe the actual behavior of any gnuplot version since before the start of the CVS repository". I have now updated the 4.5 documentation a bit. > While the subject is being raised , may I suggest a couple of possible > feature improvements in this area? > > 1) This cycling of "mode" could also be done by clicking on the status > message (click is mapped to "1" key). This would avoid switching mouse > to keyboard which can be handy. Could be done, certainly, but there is currently no code to interpret the mouse location in terms like "on the status message". > 2) As I understand it this only shows x1y1 coords. It would clearly be > useful when using x2y2 as well if those were available in this read-out. > One way to trigger this would be a click on the legend entry for a > line or on/near the relevant axis. (some thought needed to avoid > ambiguity with the new SVG toggle , pause mouse or other functions. ) Huh. For me the x2y2 axes are echoed properly if they are active in the current plot. Tested on wxt, x11, and canvas. It is true that I ignored x2y2 in the very recent svg mouse-tracking code. What terminal are you using that fails to show them? Ethan |
|
From: <pl...@pi...> - 2011-05-12 17:51:24
|
On 05/12/11 17:13, sfeam (Ethan Merritt) wrote: > On Wednesday, 11 May 2011, pl...@pi... wrote: >> Hi, >> >> I work quite a lot with time data (years and month scale) . Although >> gnuplot is very flexible in reading time data and labelling x axis the >> mouse cursor seems a bit challenged. >> >> If I have "%Y %m" , having 6.004538e+09 as mouse coord is really little >> more than useless. >> >> Since the scaling is available for axis labels why isn't this applied to >> the mouse coords? It would seem to be very simple and I can't see the >> utility in the current date in seconds. >> >> Is this just an oversight? > > Mouse coordinate readout is available in 5 or 6 different formats, > including user-specified. See "help mouseformat" and the help > message printed by typing "h" in the plot window. > In general you can cycle through the different options using the > characters "1" and "2" as hotkeys. > > Hi Ethan, thanks for that. It appears that there is not mouseformat help but I did find it under mouse. gnuplot> help mouseformat Sorry, no help for 'mouseformat' gnuplot> help mouse I had tried the h key help but had not appreciated what those variables referred to: 1 `builtin-decrement-mousemode` 2 `builtin-increment-mousemode` Maybe those terms could be more explicit, it's not clear what they refer to. The content of the various modes is equally a bit arcane and the help doesn't: >> This corresponds to the key >> bindings '1', '2', '3', '4' (see the driver's documentation) I find no detail for example in help wxt. What does this comment refer to? While the subject is being raised , may I suggest a couple of possible feature improvements in this area? 1) This cycling of "mode" could also be done by clicking on the status message (click is mapped to "1" key). This would avoid switching mouse to keyboard which can be handy. 2) As I understand it this only shows x1y1 coords. It would clearly be useful when using x2y2 as well if those were available in this read-out. One way to trigger this would be a click on the legend entry for a line or on/near the relevant axis. (some thought needed to avoid ambiguity with the new SVG toggle , pause mouse or other functions. ) Not priority issues but hopefully useful. best regards, Peter. |
|
From: sfeam (E. Merritt) <eam...@gm...> - 2011-05-12 15:14:11
|
On Wednesday, 11 May 2011, pl...@pi... wrote: > Hi, > > I work quite a lot with time data (years and month scale) . Although > gnuplot is very flexible in reading time data and labelling x axis the > mouse cursor seems a bit challenged. > > If I have "%Y %m" , having 6.004538e+09 as mouse coord is really little > more than useless. > > Since the scaling is available for axis labels why isn't this applied to > the mouse coords? It would seem to be very simple and I can't see the > utility in the current date in seconds. > > Is this just an oversight? Mouse coordinate readout is available in 5 or 6 different formats, including user-specified. See "help mouseformat" and the help message printed by typing "h" in the plot window. In general you can cycle through the different options using the characters "1" and "2" as hotkeys. |
|
From: <pl...@pi...> - 2011-05-12 06:35:42
|
Hi, I work quite a lot with time data (years and month scale) . Although gnuplot is very flexible in reading time data and labelling x axis the mouse cursor seems a bit challenged. If I have "%Y %m" , having 6.004538e+09 as mouse coord is really little more than useless. Since the scaling is available for axis labels why isn't this applied to the mouse coords? It would seem to be very simple and I can't see the utility in the current date in seconds. Is this just an oversight? regards, Peter. |
|
From: Bastian M. <bma...@we...> - 2011-05-07 19:40:51
|
A new testing binary for Windows can be found at http://www.gnuplot.info/development/binaries/ . It includes support for cairo terminals (pngcairo, pdfcairo, wxt), gd terminals (gif,png,jpeg) and lua (tikz). Bastian |
|
From: Thomas M. <mat...@ph...> - 2011-05-06 23:50:48
|
On 2011-05-06, at 3:56 PM, Daniel J Sebald wrote:
> On 05/06/2011 04:10 PM, Thomas Mattison wrote:
>>
>> 2. Gnuplot will indeed refuse to fit a constant function to a single
>> data point, or a line to two points. Yes, I hadn't done the
>> experiment. Sorry.
>>
>> However I regard item 2 as a bug not a feature. It is perfectly well
>> defined to ask the question, what is the chisquare for agreement
>> between data and model, as a function of the parameter values, for
>> these cases. There is a perfectly well defined answer to that
>> question. There is a perfectly well defined point of minimum
>> chisquare, which gives the best fit parameters. And the curvature of
>> the chisquare curve or surface gives the fit errors. Nothing
>> anomalous happens in the case where the number of data points equals
>> the number of parameters.
>
> It seems like it should.
But it doesn't, from a mathematical point of view.
> From H.B.B.'s last post, if there are only two points and the curve to fit is a line (two parameters, first order equation**) then the fit is exact in the sense that the residuals are zero...if no other information is given about the data. But that doesn't say anything about the measurement "error" in the fit. (I sort of prefer the word "uncertainty", but that's just me.) We can only glean information about the uncertainty of the fit (i.e., the statistics of the original data) if there are additional data points beyond the number of parameters.
You are right that there need to be more measurements than parameters in order to use the chisquare/degree-of-freedom to infer something about the errors of the fit INPUTS from the quality of the fit.
But in many cases, the user ALREADY KNOWS what the input data measurement errors are. He knows the precision of the markings on his ruler, or how many decimal places his voltmeter has, or square-root-of-N for the bin contents of a histogram, etc [I am not discussing "systematic" errors here, like is the calibration of your voltmeter correct].
In those cases, the "raw" fit errors are the appropriate ones, and the "rescaled" errors are less appropriate.
>
> OK, so I'm beginning to see there is this vital piece of information about the quality of fit, reflecting the underlying statistics or the uncertainty in the data (i.e., how noisy it is). That's what is meant by error, right?
The way I would phrase it is, the noise in the input data propagates to noise in the fit parameters. The "standard errors" of the fit parameters are the standard deviation of the input data (assumed to be known) propagated to the standard deviation of the fit parameters.
>
>> I hope we can all agree that when fitting a line to two data points
>> that have errors, the slope and intercept DO have errors.
>
> Yes, but I don't see how one can estimate the size of that error without additional data beyond two data points, unless some information is known apriori.
>
Sure, but when the user has supplied meaningful y-errors on his data, we have precisely the a-priori informaton that your intuition is telling you is needed.
> But I'm surprised that no well defined, commonly understood terminology has developed for what seems like an important idea in the numerical analysis community. Bastian points out that:
>
> * Minuit does not scale
> * Origin offers both options, the default being version dependent
> * SAS and Mathematica default to scaling of errors
>
> That inconsistency is the problem.
To me, that looks like most programs give you a choice. Mathematica may default to scaling, but it looks like it's optional. I don't have much knowledge about SAS, but I would expect a big package like that to have an option buried somewhere.
Minuit comes from the physics community, nuclear and particle physics in particular. In physics, we tend to have pretty good a-priori estimates for the errors of the measurements going into a fit. We use the chisquare/DOF as a measure of the (statistical) fit quality. Physicists would seldom rescale fit parameter errors by sqrt(chisq/DOF). And since Minuit is normally used in an expert-programming context instead of an ignorant-GUI-user context, the authors would expect the user to be able multiply by sqrt(chisq/DOF) himself.
In bioscience, and even more so in medical and social science, there's often next to nothing known about the errors of the input measurements a-priori. In such cases, there's not much the user can do except set all the errors equal (or not tell gnuplot any errors, which will have the same result). Then the "raw" errors from the fit aren't very meaningful, but the "rescaled" errors have at least some utility. So defaulting to rescaled makes some sence.
My guess would be that there are more gnuplot fit users who don't know much about their input data errors than those who know a great deal about their input data errors. So if I were forced to make a choice of only one error to present, I'd probably present the rescaled error. It will be the "right thing" for users who don't give errors or give uniform and arbitrary errors. And for the people who supply accurate errors, if the model really reflects the data, then chisq/DOF will be close to 1, and the "rescaled" errors will not be very different, very often, from the "raw" errors.
But why not give both?
Cheers
Prof. Thomas Mattison Hennings 276
University of British Columbia Dept. of Physics and Astronomy
6224 Agricultural Road Vancouver BC V6T 1Z1 CANADA
mat...@ph... phone: 604-822-9690 fax:604-822-5324
|
|
From: Daniel J S. <dan...@ie...> - 2011-05-06 22:56:48
|
On 05/06/2011 04:10 PM, Thomas Mattison wrote: > I stand corrected about two things: > > 1. Gnuplot fits have always done the "right" thing and divided > chisquare by degrees of freedom rather than measurements. Sorry, I > haven't looked at what gets printed out in a while; I mostly use my > own fitting programs these days. > > 2. Gnuplot will indeed refuse to fit a constant function to a single > data point, or a line to two points. Yes, I hadn't done the > experiment. Sorry. > > However I regard item 2 as a bug not a feature. It is perfectly well > defined to ask the question, what is the chisquare for agreement > between data and model, as a function of the parameter values, for > these cases. There is a perfectly well defined answer to that > question. There is a perfectly well defined point of minimum > chisquare, which gives the best fit parameters. And the curvature of > the chisquare curve or surface gives the fit errors. Nothing > anomalous happens in the case where the number of data points equals > the number of parameters. It seems like it should. From H.B.B.'s last post, if there are only two points and the curve to fit is a line (two parameters, first order equation**) then the fit is exact in the sense that the residuals are zero...if no other information is given about the data. But that doesn't say anything about the measurement "error" in the fit. (I sort of prefer the word "uncertainty", but that's just me.) We can only glean information about the uncertainty of the fit (i.e., the statistics of the original data) if there are additional data points beyond the number of parameters. ** Polynomial order of the fitting curve is important. For example, exp() has infinite polynomial order, which has ramifications on uniqueness. > I agree that in those two particular cases, it's not necessary to "do > a fit" to find the parameters; it can easily be done by hand. But > mathematically the fit is still well-defined. OK, so I'm beginning to see there is this vital piece of information about the quality of fit, reflecting the underlying statistics or the uncertainty in the data (i.e., how noisy it is). That's what is meant by error, right? In this discussion http://www.erikburd.org/projects/pitfall/evaluate.html is a column for "standard error". Is that what is being referred to, the standard error typically used in statistical analysis? (There are the underlying actual measurement errors, but we can't know what those are because there is uncertainty in the fit.) > I hope we can all agree that when fitting a line to two data points > that have errors, the slope and intercept DO have errors. Yes, but I don't see how one can estimate the size of that error without additional data beyond two data points, unless some information is known apriori. > Is it really so much to ask to have both raw and rescaled errors > printed? I'm fine with that idea. But I'm surprised that no well defined, commonly understood terminology has developed for what seems like an important idea in the numerical analysis community. Bastian points out that: * Minuit does not scale * Origin offers both options, the default being version dependent * SAS and Mathematica default to scaling of errors That inconsistency is the problem. Dan |
|
From: Thomas M. <mat...@ph...> - 2011-05-06 21:10:45
|
I stand corrected about two things:
1. Gnuplot fits have always done the "right" thing and divided chisquare by degrees of freedom rather than measurements. Sorry, I haven't looked at what gets printed out in a while; I mostly use my own fitting programs these days.
2. Gnuplot will indeed refuse to fit a constant function to a single data point, or a line to two points. Yes, I hadn't done the experiment. Sorry.
However I regard item 2 as a bug not a feature. It is perfectly well defined to ask the question, what is the chisquare for agreement between data and model, as a function of the parameter values, for these cases. There is a perfectly well defined answer to that question. There is a perfectly well defined point of minimum chisquare, which gives the best fit parameters. And the curvature of the chisquare curve or surface gives the fit errors. Nothing anomalous happens in the case where the number of data points equals the number of parameters.
I agree that in those two particular cases, it's not necessary to "do a fit" to find the parameters; it can easily be done by hand. But mathematically the fit is still well-defined.
I hope we can all agree that when fitting a line to two data points that have errors, the slope and intercept DO have errors. But calculating them by hand is not as trivial as calculating the slope and intercept by hand. The easiest way to do that in practice is to run a fitting program. The "raw" errors from the fit give the answer.
But gnuplot refuses to do that. And the only reason I can see for the refusal is that the "rescaled" errors wouldn't make sense in that case (while the raw errors would make sense, and are useful).
Let's say I have several equal-size data sets, described by the same model, with the same amount of random noise in each. If I fit each data set, the parameters will vary (within the errors), but they all have equal amounts of information, so they SHOULD all give the same fit errors. The "raw" fit errors WILL be the (neglecting effects from nonlinearities which are usually small). But the "rescaled" fit errors will vary randomly, depending on the chisquare of the fit. For a case where someone has worked hard to make the input errors meaningful, it's wrong for the parameter error output to depend on randomness in the data.
Is it really so much to ask to have both raw and rescaled errors printed?
Cheers
Prof. Thomas Mattison Hennings 276
University of British Columbia Dept. of Physics and Astronomy
6224 Agricultural Road Vancouver BC V6T 1Z1 CANADA
mat...@ph... phone: 604-822-9690 fax:604-822-5324
|
|
From: Hans-Bernhard B. <HBB...@t-...> - 2011-05-06 20:03:50
|
On 06.05.2011 19:48, Thomas Mattison wrote:
> Consider the case where the "function" to be fit is just a constant:
> we think the the y-value of all data points should be the same,
> independent of x-value. Let's call this constant C. Surely we want
> gnuplot to be able to do this "fit" and give a meaningful error on
> C.
Yes.
> Then consider the special case where there is only a single data
> point: an x-value, a y-value, and a y-error. The fit result should
> C=y, and the error should be sigma-C = sigma-y.
It shouldn't. The problem here is that by going down to a single data
point, you've constructed a problem with zero degrees of
freedom. That's not a fit --- it's just an equation to be evaluated.
Using fit for that is tool abuse.
> The fit should be "perfect" with a chisquare of zero. The chisquare
> divided by the number of measurements is still zero.
But chisq is never divided by the number of measurements. It's divided
by the number of degrees of freedom, i.e. (number of measurements) -
(number of parameters). In the case at hand that's zero, STDFIT would
be zero divided by zero.
> Internally, gnuplot's fit algorithm knows a number equivalent to
> sigma-C = sigma-y,
No, it doesn't, because it never even starts to work in that case:
gnuplot> fit a '-' u 1:2:3 via a
input data ('e' ends) > 0 10 2
input data ('e' ends) > e
Read 1 points
No data to fit
> but the reported error is that multiplied by chisquare/measurement,
> so the result is zero.
No it's not. See above.
> It is totally unreasonable to claim that "fitting" the single data
> point gives C with infinite precision, when the data point has a
> finite error.
It's totally unreasonable to use fit on a single data point, period.
> One can make a similar argument about the case of fitting a line (two
> parameters, slope and intercept) to two x-y-yerror data points.
Same problem. Zero degrees of freedom means 'fit' is the wrong tool for
the job.
> Gnuplot's fit will go exactly through both data points, with zero
> chisquare per measurement, so the reported errors on the slope and
> intercept will both be zero.
No they won't. Because gnuplot won't report _any_ errors in that case.
> But clearly there is a range of slopes
> and range of intercepts that are "within one sigma" of both data
> points. And internally gnuplot's fit algorithm knows those errors on
> slope and intercept, before it multiplies them by zero.
No, it doesn't.
> I would make one small additional change: chisquare should be divided
> by "degrees of freedom" == (measurements minus parameters) not just
> measurements. Do this both for reporting STDFIT and rescaling of
> errors.
What made you believe that's not already what's happening?
> user that the rescaled errors are meaningless. This is much more
> sensible than what gnuplot does now, which is report the errors are
> zero.
Your view of the world might profit from an actual experiment.
|
|
From: Hans-Bernhard B. <HBB...@t-...> - 2011-05-06 19:45:18
|
On 06.05.2011 05:08, Daniel J Sebald wrote: > On 05/05/2011 04:10 PM, Bastian Märkisch wrote: > One thing I notice is that Hans-Bernhard uses the term "residuals". Yes, but only for what's always called that: the difference between fitted model and data. > Now, "residuals" is a fairly common definition in fitting. Maybe it > would have been better in the first place to use "_res" extensions to > variable names for the unscaled errors and "_err" for the scaled errors I rather much doubt that. The term residual is not applicable to parameter errors in any meaningful way. > The argument is made in the post that one is derived from the other with > a simple scaling; That argument would be wrong. Parameter errors scale in parallel with the data errors (or weights, if you prefer), not with the residiuals. > avoid ambiguity. If "error" has some ambiguity in the field, whereas > "residual" is much less ambiguous, then go with the latter. "Residual" is unambiguous primarily because it is use for exactly one purpose. Calling something else by the same name would only break that unambiguity. That wouldn't be particularly helpful. |
|
From: Daniel J S. <dan...@ie...> - 2011-05-06 17:55:51
|
On 05/05/2011 06:30 PM, Hans-Bernhard Bröker wrote: > On 05.05.2011 10:52, Daniel J Sebald wrote: [snip] >> As for the original bug report, unless this is something obvious, >> perhaps there is a way to illustrate the error with a test case, to >> ensure the fit is solved correctly. > > The demo is dead simple. Pick any fit from the demos or wherever, and > repeat it with the data errors multiplied by a fixed factor, i.e. replace > > fit f(x) 'foo.dat' u 1:2:3 via ... > > by > > fit f(x) 'foo.dat' u 1:2:($3*20) via ... > > 'fit' will report the same data errors, both in the printed output and > in the saved *_err variables. Only the chisq and STDFIT will have > shrinked by a factor of 20. > > People thinking I made a bad decision here say that the errors on the > parameters should become 20 times as large in the second case. Well, this is certainly a valid approach. Often some statistical or algebraic quantity is independent of scale to reflect is quality or fundamental nature. Could we create an additional set of variables "_res" that is scaled as others might want (i.e., residuals)? Dan |
|
From: Thomas M. <mat...@ph...> - 2011-05-06 17:48:58
|
Hi
Consider the case where the "function" to be fit is just a constant: we think the the y-value of all data points should be the same, independent of x-value. Let's call this constant C. Surely we want gnuplot to be able to do this "fit" and give a meaningful error on C.
Then consider the special case where there is only a single data point: an x-value, a y-value, and a y-error. The fit result should C=y, and the error should be sigma-C = sigma-y. The fit should be "perfect" with a chisquare of zero. The chisquare divided by the number of measurements is still zero.
Internally, gnuplot's fit algorithm knows a number equivalent to sigma-C = sigma-y, but the reported error is that multiplied by chisquare/measurement, so the result is zero.
It is totally unreasonable to claim that "fitting" the single data point gives C with infinite precision, when the data point has a finite error.
One can make a similar argument about the case of fitting a line (two parameters, slope and intercept) to two x-y-yerror data points. Gnuplot's fit will go exactly through both data points, with zero chisquare per measurement, so the reported errors on the slope and intercept will both be zero. But clearly there is a range of slopes and range of intercepts that are "within one sigma" of both data points. And internally gnuplot's fit algorithm knows those errors on slope and intercept, before it multiplies them by zero.
Every other fitting program that I have ever used, which takes into account y-errors, just reports the "internal" error and doesn't divide it by chisquare/measurement the way gnuplot does.
If the y-errors are known and are "correct", the points will scatter around the fit curve such that the residual (measurement minus curve) divided by the error has a gaussian distribution with mean of zero and sigma of 1.0. If the "real" y-errors are twice as large as the "assumed" y-errors, the gaussian will have sigma of about 2.0, and the chisquare/measurement will be about 4.0.
However, the user is not required to provide y-errors to gnuplot. In that case, the errors of the fit parameters are more problematic. Gnuplot does what every other no-y-errors fitting program does: treat all the y-errors as being 1.0. The "internal" error calculation is still done, but one does not expect it to really describe the errors of the parameters. And one does not expect the "chisquare" per measurement to be approximately 1.0.
But we can use the "chisquare" per measurement to estimate the y-error, assuming that it's the same for all measurements. The result is that y-error is just sqrt(chisquare/measurement).
If we then repeated the fit using those y-errors, we do expect that the "internal" parameter errors will describe our knowledge of the parameters.
In fact, it's not necessary to actually repeat the fit. The parameter values would come out exactly the same, and the parameter errors would just be multiplied by sqrt(chisquare/measurement).
This is in fact what gnuplot's fit algorithm does. And for the case where no user errors are provided, I think it's (almost) the most sensible thing to do.
So what should be done? If it were me, I'd report both the "raw" and the "rescaled" parameter errors on the console. The documentation should state that if the user has provided errors, the "raw" errors are the most meaningful, and if the user has not provided errors, the "rescaled" errors are the most meaningful. This requires essentially no algorithm changes, just some print statement changes.
Perhaps at the end of the fit, one could print a recommendation to the user on which error to use, based on whether or not he provided errors, and on the chisquare per measurement. If no errors were provided, only the rescaled errors are meaningful. If errors were provided and the chisquare per measurement is reasonable (less than 2 perhaps), the raw errors are meaningful. If errors were provided and the chisquare per measurement is not reasonable, then the raw errors are questionable and the rescaled errors might be more meaningful.
I would make one small additional change: chisquare should be divided by "degrees of freedom" == (measurements minus parameters) not just measurements. Do this both for reporting STDFIT and rescaling of errors. Dividing by degrees of freedom instead of measurements is totally standard statistical practice.
In the one-parameter fit to one data point and line-fit to two data points examples I give above, measurements-parameters is zero. Ideally the chisquare is also zero, so ideally the rescaling factor is 0/0=NAN, which gives the "correct" result that the rescaled errors are meaningless (in these cases). In practice the chisquare won't be exactly zero, so the rescaling factor will be INF, still telling the user that the rescaled errors are meaningless. This is much more sensible than what gnuplot does now, which is report the errors are zero.
If we were fitting a line to 3 points instead of 2, with errors supplied, we expect the average raw chisquare to be 1, not 3. If errors are not supplied, we expect the raw "chisquare" not to be 3*yerror^2, but only 1*yerror^2. Dividing by (measurements - parameters) gives a significantly different (and more correct) estimate of the yerrors, and thus of the parameter errors.
Cheers
Prof. Thomas Mattison Hennings 276
University of British Columbia Dept. of Physics and Astronomy
6224 Agricultural Road Vancouver BC V6T 1Z1 CANADA
mat...@ph... phone: 604-822-9690 fax:604-822-5324
|
|
From: <pl...@pi...> - 2011-05-06 13:55:34
|
On 05/05/11 17:40, sfeam (Ethan Merritt) wrote:
> On Thursday, 05 May 2011, pl...@pi... wrote:
>> Any reason that is not in the install instructions in INSTALL?
>
> The CVS source tree contains a file INSTALL for use with the
> eventual distributed release package. The instructions in
> INSTALL are correct for the release package. But if you are
> building directly from the CVS source tree you need to first run
> the files through it through autoconf.
> There is a script "./prepare" that does this for you.
>
> Ethan
>
Allin Cottrell wrote:
You need to read README.1ST ;-)
Hi,
Allin is of course correct. Thanks for pointing to that file:
If your source is from the CVS repository rather from a release package,
the order of commands is ./prepare; ./configure; make
So the point is that INSTALL, despite it's length, does not tell you how
to install in all cases. There are two files giving install notes with a
paragraph headed "Installation from sources
", the two with differing information.
Is there any pressing reason for this simple phrase about how to install
not being in the file called INSTALL?
Why is there information in README.1ST: 'Installation from sources' that
is not in INSTALL: 'Installation from sources' ?
Surely the two should be syncronised or the para from README.1ST should
be merged into INSTALL.
thanks for the replies.
regards, Peter.
|