You can subscribe to this list here.
| 2001 |
Jan
|
Feb
(1) |
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
|
Sep
|
Oct
|
Nov
|
Dec
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2002 |
Jan
(1) |
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
|
Nov
(1) |
Dec
|
| 2003 |
Jan
|
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
(83) |
Nov
(57) |
Dec
(111) |
| 2004 |
Jan
(38) |
Feb
(121) |
Mar
(107) |
Apr
(241) |
May
(102) |
Jun
(190) |
Jul
(239) |
Aug
(158) |
Sep
(184) |
Oct
(193) |
Nov
(47) |
Dec
(68) |
| 2005 |
Jan
(190) |
Feb
(105) |
Mar
(99) |
Apr
(65) |
May
(92) |
Jun
(250) |
Jul
(197) |
Aug
(128) |
Sep
(101) |
Oct
(183) |
Nov
(186) |
Dec
(42) |
| 2006 |
Jan
(102) |
Feb
(122) |
Mar
(154) |
Apr
(196) |
May
(181) |
Jun
(281) |
Jul
(310) |
Aug
(198) |
Sep
(145) |
Oct
(188) |
Nov
(134) |
Dec
(90) |
| 2007 |
Jan
(134) |
Feb
(181) |
Mar
(157) |
Apr
(57) |
May
(81) |
Jun
(204) |
Jul
(60) |
Aug
(37) |
Sep
(17) |
Oct
(90) |
Nov
(122) |
Dec
(72) |
| 2008 |
Jan
(130) |
Feb
(108) |
Mar
(160) |
Apr
(38) |
May
(83) |
Jun
(42) |
Jul
(75) |
Aug
(16) |
Sep
(71) |
Oct
(57) |
Nov
(59) |
Dec
(152) |
| 2009 |
Jan
(73) |
Feb
(213) |
Mar
(67) |
Apr
(40) |
May
(46) |
Jun
(82) |
Jul
(73) |
Aug
(57) |
Sep
(108) |
Oct
(36) |
Nov
(153) |
Dec
(77) |
| 2010 |
Jan
(42) |
Feb
(171) |
Mar
(150) |
Apr
(6) |
May
(22) |
Jun
(34) |
Jul
(31) |
Aug
(38) |
Sep
(32) |
Oct
(59) |
Nov
(13) |
Dec
(62) |
| 2011 |
Jan
(114) |
Feb
(139) |
Mar
(126) |
Apr
(51) |
May
(53) |
Jun
(29) |
Jul
(41) |
Aug
(29) |
Sep
(35) |
Oct
(87) |
Nov
(42) |
Dec
(20) |
| 2012 |
Jan
(111) |
Feb
(66) |
Mar
(35) |
Apr
(59) |
May
(71) |
Jun
(32) |
Jul
(11) |
Aug
(48) |
Sep
(60) |
Oct
(87) |
Nov
(16) |
Dec
(38) |
| 2013 |
Jan
(5) |
Feb
(19) |
Mar
(41) |
Apr
(47) |
May
(14) |
Jun
(32) |
Jul
(18) |
Aug
(68) |
Sep
(9) |
Oct
(42) |
Nov
(12) |
Dec
(10) |
| 2014 |
Jan
(14) |
Feb
(139) |
Mar
(137) |
Apr
(66) |
May
(72) |
Jun
(142) |
Jul
(70) |
Aug
(31) |
Sep
(39) |
Oct
(98) |
Nov
(133) |
Dec
(44) |
| 2015 |
Jan
(70) |
Feb
(27) |
Mar
(36) |
Apr
(11) |
May
(15) |
Jun
(70) |
Jul
(30) |
Aug
(63) |
Sep
(18) |
Oct
(15) |
Nov
(42) |
Dec
(29) |
| 2016 |
Jan
(37) |
Feb
(48) |
Mar
(59) |
Apr
(28) |
May
(30) |
Jun
(43) |
Jul
(47) |
Aug
(14) |
Sep
(21) |
Oct
(26) |
Nov
(10) |
Dec
(2) |
| 2017 |
Jan
(26) |
Feb
(27) |
Mar
(44) |
Apr
(11) |
May
(32) |
Jun
(28) |
Jul
(75) |
Aug
(45) |
Sep
(35) |
Oct
(285) |
Nov
(99) |
Dec
(16) |
| 2018 |
Jan
(8) |
Feb
(8) |
Mar
(42) |
Apr
(35) |
May
(23) |
Jun
(12) |
Jul
(16) |
Aug
(11) |
Sep
(8) |
Oct
(16) |
Nov
(5) |
Dec
(8) |
| 2019 |
Jan
(9) |
Feb
(28) |
Mar
(4) |
Apr
(10) |
May
(7) |
Jun
(4) |
Jul
(4) |
Aug
|
Sep
(4) |
Oct
|
Nov
(23) |
Dec
(3) |
| 2020 |
Jan
(19) |
Feb
(3) |
Mar
(22) |
Apr
(17) |
May
(10) |
Jun
(69) |
Jul
(18) |
Aug
(23) |
Sep
(25) |
Oct
(11) |
Nov
(20) |
Dec
(9) |
| 2021 |
Jan
(1) |
Feb
(7) |
Mar
(9) |
Apr
|
May
(1) |
Jun
(8) |
Jul
(6) |
Aug
(8) |
Sep
(7) |
Oct
|
Nov
(2) |
Dec
(23) |
| 2022 |
Jan
(23) |
Feb
(9) |
Mar
(9) |
Apr
|
May
(8) |
Jun
(1) |
Jul
(6) |
Aug
(8) |
Sep
(30) |
Oct
(5) |
Nov
(4) |
Dec
(6) |
| 2023 |
Jan
(2) |
Feb
(5) |
Mar
(7) |
Apr
(3) |
May
(8) |
Jun
(45) |
Jul
(8) |
Aug
|
Sep
(2) |
Oct
(14) |
Nov
(7) |
Dec
(2) |
| 2024 |
Jan
(4) |
Feb
(4) |
Mar
|
Apr
(7) |
May
(2) |
Jun
(1) |
Jul
|
Aug
(5) |
Sep
|
Oct
|
Nov
(4) |
Dec
(14) |
| 2025 |
Jan
(22) |
Feb
(6) |
Mar
(5) |
Apr
(14) |
May
(6) |
Jun
(11) |
Jul
(19) |
Aug
|
Sep
(17) |
Oct
(1) |
Nov
(2) |
Dec
(18) |
| 2026 |
Jan
|
Feb
|
Mar
(5) |
Apr
|
May
(2) |
Jun
(1) |
Jul
(6) |
Aug
(1) |
Sep
|
Oct
|
Nov
|
Dec
|
|
From: Eric S. R. <es...@th...> - 2017-10-21 22:42:08
|
Daniel J Sebald <dan...@ie...>: > However, I don't understand the requirement that the date be unique. The > only changes to be made for the changeset are Author and Email. The more > detailed information about the SHA, commiter, etc. remains the same. Why > the requirement? It's so each commit can be identified by a unique action stamp based on its author address and authorship time. Otherwise it's difficult for reposurgeon to have a unique way to refer to commits, which makes it difficult to write surgical commands. Why the athor date and not the committer date, which we always know to 1-second precision? Ah, but the author stamp doesn't change when author-attributed patches are replayed onto a repository with git am, while the commit date does. The author date is a property of the patch, the committer date an artifact Unfortunately, author attributions mined from ChangeLogs only have time specified as a a date. In an attempt to minimize the worst-case distance from the actual time of authorship, I add on a time part of 12:00:00Z. Kind of doomed since we don't know the submitter's time zone, but I had to pick something and noon UTC will work pretty well for Europe and the U.S. This gives rise to another problem - lots of noon timestamps colliding with each other (making for non-unique action stamps if one author has multiple commits on the same day). One thing I think I know is that order of attributions in a ChangeLog is usuall the time order, so I add a tick to the timetamps to separate them. Maybe it would be better to copy the committer time if it's on the same day as the ChangeLog entry. We still have to deal with the other case, though. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Eric S. R. <es...@th...> - 2017-10-21 21:34:12
|
"Bastian Märkisch" <bma...@we...>: > No quite the end of the story, though. With that change I get 135 > "no time stamp matching" errors. That's odd. I can't find anywjhere my code emits that message. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Eric S. R. <es...@th...> - 2017-10-21 21:31:49
|
"Bastian Märkisch" <bma...@we...>: > I think Daniel's conclusion is correct. The script does not break out of the > loop for non-author lines and hence may read in author lines below the changeline. > Adding > if n > changeline: > break > at the top of the loop eliminates that. > > Please find a test case attached which demonstrates the problem. I've uploaded a new test tarball that you can get with wget http://www.catb.org/~esr/gnuplot-conversion.tar.gz Please check it against your edge case. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Eric S. R. <es...@th...> - 2017-10-21 21:30:54
|
Daniel J Sebald <dan...@ie...>:
> From step 3, I think you meant to place the test of 'n' and 'changeline' at
> the front of loop and 'break' (on >) instead. I.e.,
>
> if n > changeline:
> break
> if ',' in line:
> # Multiple attributions...ignore for now
> continue
> # Deal with some address masking
> line = line.replace(" <at> ", "@")
> space = line.find(" ")
> if space < 0:
> continue
> ETC.
>
> It should break if "n > changeline", not if "n >= changeline", because we
> want to include the scenario of author info at the first changed line (which
> is the most typical scenario).
That change didn't work - broke my regression test for two other cases. But
I found an easier-to-understand change that did. The liftlog regression test
now verifies three cases -- new entry at to of file, new entry within the file,
and addition of text to an entry within the file.
> While on the subject, I see you've include the "if only ChangeLog changes,
> ignore" heuristic. What about the case of the change in the ChangeLog being
> a subtraction rather than an addition. I left that out because all the
> scenarios I imagined that happening were a situation we wanted to ignore any
> authorship change and just let the committer have authorship (e.g.,
> wholesale swap of ChangeLog.0 to ChangeLog.1, some typo in the authorship
> line was corrected). I'm not sure your code differentiates between the two
> (it looks to be searching just for change), but really this is a very low
> likelihood of occurrence so maybe it isn't worth toiling over that.
I don't think so. Unless a real case smacks ua in the case, anyway.
I've uploaded a new test tarball that you can get with
wget http://www.catb.org/~esr/gnuplot-conversion.tar.gz
Please check it against your edge cases.
--
<a href="http://www.catb.org/~esr/">Eric S. Raymond</a>
My work is funded by the Internet Civil Engineering Institute: https://icei.org
Please visit their site and donate: the civilization you save might be your own.
|
|
From: Daniel J S. <dan...@ie...> - 2017-10-21 20:53:10
|
On 10/21/2017 03:31 PM, "Bastian Märkisch" wrote:
>> Gesendet: Samstag, 21. Oktober 2017 um 22:08 Uhr
>> Von: "Daniel J Sebald" <dan...@ie...>
>> An: "Bastian Märkisch" <bma...@we...>
>> Cc: es...@th..., gnu...@li...
>> Betreff: Re: Aw: Re: News spin of repository conversion
>>
>> On 10/21/2017 02:53 PM, "Bastian Märkisch" wrote:
>>>>
>>>> Hence, what is really happening above is a search for the first author
>>>> line that comes at or after 'changeline'. That's not correct.
>>>>
>>>> From step 3, I think you meant to place the test of 'n' and
>>>> 'changeline' at the front of loop and 'break' (on >) instead. I.e.,
>>>>
>>>> if n > changeline:
>>>> break
>>>> if ',' in line:
>>>> # Multiple attributions...ignore for now
>>>> continue
>>>> # Deal with some address masking
>>>> line = line.replace(" <at> ", "@")
>>>> space = line.find(" ")
>>>> if space < 0:
>>>> continue
>>>> ETC.
>>>>
>>>> It should break if "n > changeline", not if "n >= changeline", because
>>>> we want to include the scenario of author info at the first changed line
>>>> (which is the most typical scenario).
>>>
>>> No quite the end of the story, though. With that change I get 135
>>> "no time stamp matching" errors.
>>
>> I see the code that checks the date and makes sure that the date is unique:
>>
>> if fdate in generated:
>> continue
>>
>> However, I don't understand the requirement that the date be unique.
>> The only changes to be made for the changeset are Author and Email. The
>> more detailed information about the SHA, commiter, etc. remains the
>> same. Why the requirement?
>>
>> Dan
>
> No idea. But the messages about non-matched tags come from the LONGLINES
> file. Should have seen that in the log...
>
> I can still identifiy a few missing and false attributions. I guess those
> could be included easily with the help of an additional input file, right?
If they look like something that should fall under the
first-author-line-previous-to-change rule, then we should address them.
But if due to some unique peculiar edit, then yes put those in a file.
I'm not sure it pays to make the code have too many conditionals to
catch corner cases, plus it becomes less generic the more unique
circumstances addressed.
One thing I wondered about--and I'm just in the process of updating my
utility to address this--is if somehow the first line changed is
something other than the typical author info, e.g.,
* src/win/wd2d.cpp: Resize the swap chain buffers instead of
recreating the swap chain when the window size changes.
+
+2017-07-24 Ethan A Merritt <merritt@u.washington.edu>
+
+ * src/win/wgraph.c term/win.trm: Default to Direct2D backend.
I can easily address that in the utility by starting not from the first
'+' change, but from the last '+' change of the first block of '+'
changes. However, I'm not sure that can be done with Eric's python
algorithm because one needs to differentiate between the addition '+',
subtraction '-'. The python approach only differentiates between
modified line ('+' or '-') and non-modified line (' '). The '+' and '-'
come from the diff utility, which is surprisingly sophisticated.
Dan
|
|
From: Bastian M. <bma...@we...> - 2017-10-21 20:31:22
|
> Gesendet: Samstag, 21. Oktober 2017 um 22:08 Uhr
> Von: "Daniel J Sebald" <dan...@ie...>
> An: "Bastian Märkisch" <bma...@we...>
> Cc: es...@th..., gnu...@li...
> Betreff: Re: Aw: Re: News spin of repository conversion
>
> On 10/21/2017 02:53 PM, "Bastian Märkisch" wrote:
> >>
> >> Hence, what is really happening above is a search for the first author
> >> line that comes at or after 'changeline'. That's not correct.
> >>
> >> From step 3, I think you meant to place the test of 'n' and
> >> 'changeline' at the front of loop and 'break' (on >) instead. I.e.,
> >>
> >> if n > changeline:
> >> break
> >> if ',' in line:
> >> # Multiple attributions...ignore for now
> >> continue
> >> # Deal with some address masking
> >> line = line.replace(" <at> ", "@")
> >> space = line.find(" ")
> >> if space < 0:
> >> continue
> >> ETC.
> >>
> >> It should break if "n > changeline", not if "n >= changeline", because
> >> we want to include the scenario of author info at the first changed line
> >> (which is the most typical scenario).
> >
> > No quite the end of the story, though. With that change I get 135
> > "no time stamp matching" errors.
>
> I see the code that checks the date and makes sure that the date is unique:
>
> if fdate in generated:
> continue
>
> However, I don't understand the requirement that the date be unique.
> The only changes to be made for the changeset are Author and Email. The
> more detailed information about the SHA, commiter, etc. remains the
> same. Why the requirement?
>
> Dan
No idea. But the messages about non-matched tags come from the LONGLINES
file. Should have seen that in the log...
I can still identifiy a few missing and false attributions. I guess those
could be included easily with the help of an additional input file, right?
Bastian
|
|
From: Daniel J S. <dan...@ie...> - 2017-10-21 20:08:55
|
On 10/21/2017 02:53 PM, "Bastian Märkisch" wrote:
>>
>> Hence, what is really happening above is a search for the first author
>> line that comes at or after 'changeline'. That's not correct.
>>
>> From step 3, I think you meant to place the test of 'n' and
>> 'changeline' at the front of loop and 'break' (on >) instead. I.e.,
>>
>> if n > changeline:
>> break
>> if ',' in line:
>> # Multiple attributions...ignore for now
>> continue
>> # Deal with some address masking
>> line = line.replace(" <at> ", "@")
>> space = line.find(" ")
>> if space < 0:
>> continue
>> ETC.
>>
>> It should break if "n > changeline", not if "n >= changeline", because
>> we want to include the scenario of author info at the first changed line
>> (which is the most typical scenario).
>
> No quite the end of the story, though. With that change I get 135
> "no time stamp matching" errors.
I see the code that checks the date and makes sure that the date is unique:
if fdate in generated:
continue
However, I don't understand the requirement that the date be unique.
The only changes to be made for the changeset are Author and Email. The
more detailed information about the SHA, commiter, etc. remains the
same. Why the requirement?
Dan
|
|
From: Bastian M. <bma...@we...> - 2017-10-21 19:53:47
|
>
> Hence, what is really happening above is a search for the first author
> line that comes at or after 'changeline'. That's not correct.
>
> From step 3, I think you meant to place the test of 'n' and
> 'changeline' at the front of loop and 'break' (on >) instead. I.e.,
>
> if n > changeline:
> break
> if ',' in line:
> # Multiple attributions...ignore for now
> continue
> # Deal with some address masking
> line = line.replace(" <at> ", "@")
> space = line.find(" ")
> if space < 0:
> continue
> ETC.
>
> It should break if "n > changeline", not if "n >= changeline", because
> we want to include the scenario of author info at the first changed line
> (which is the most typical scenario).
No quite the end of the story, though. With that change I get 135
"no time stamp matching" errors.
Bastian
|
|
From: Daniel J S. <dan...@ie...> - 2017-10-21 15:55:31
|
On 10/21/2017 06:42 AM, Eric S. Raymond wrote:
> Daniel J Sebald <dan...@ie...>:
>> I'd say no manual correction for that. The pool to search through is all
>> entries, 6000+. I was thinking the only manual adjustment might be the
>> substitution of "empty log message" as one of the last steps. That's only
>> 300 entries, much more reasonable to walk through.
>
> Agreed.
>
> You might want to take alook at the logic, though. I *think* my algorithm
> is equivalent to yours, but I could be wrong.
Close, but I believe there is one small detail different. I see now
this could have been done as just a single loop--hindsight--but that's
sort of too much conditional code and it's easier to understand as two
loops.
> For each changeset in the repository conversion that includes a ChangeLog blob:
>
> 1. Fetch the ancestral version of ChangeLog, if it has one.
> Compare leading lines until you hit a pair that doesn't match.
> That line number beciomes the value of 'changeline'. If you
> find no ancestor, let changeline be 0.
>
> 2. Walk through the (newer) Changelog line file. Parse each attribution
> time for date and email address, soeing them in variables 'date' and 'addr'.
>
> 3. Break out of that loop when (a) the line you just parsed is at or
> after changeline, or (b) you reach EOF.
You were thinking about this correctly, but the code is a subtle
difference in the fact that it "continues" out of the loop search when
the line in question does not match author info. By using such a
construct, the test for n < changeline does not happen unless the line
in question meets th criteria of being an author line, i.e., notice in
the following that the change in flow is always a "continue".
if ',' in line:
# Multiple attributions...ignore for now
continue
# Deal with some address masking
line = line.replace(" <at> ", "@")
space = line.find(" ")
if space < 0:
continue
date = line[:space]
# Avoid fatal error in the attribution builder
try:
time.strptime(date, "%Y-%m-%d")
except ValueError:
continue
addr = line[space+1:]
if len(date) < 27:
# Try to never assign the same generated
# timestamp. This may help later on when
# we need action stamps to be unique, e.g.
# for mailbox_in.
for offset in range(60):
fdate = date + "T12:00:%02dZ" % offset
if fdate in generated:
continue
generated.add(fdate)
date = fdate
break
else:
raise Recoverable("reposurgeon: can't uniquify ChangeLog
date.")
# Placement of this statement is very important.
# The effect we want is for all lines before the
# first changed one to be parsed for date and
# addr, but for only the *last* date and addr
# to be kept/ This handles the case where the
# changes secrtion begins inside an entry (after
# its header line) as well as the obviuos case
# where the changed section *begins* with an
# attribution line.
if n < changeline:
continue
Hence, what is really happening above is a search for the first author
line that comes at or after 'changeline'. That's not correct.
From step 3, I think you meant to place the test of 'n' and
'changeline' at the front of loop and 'break' (on >) instead. I.e.,
if n > changeline:
break
if ',' in line:
# Multiple attributions...ignore for now
continue
# Deal with some address masking
line = line.replace(" <at> ", "@")
space = line.find(" ")
if space < 0:
continue
ETC.
It should break if "n > changeline", not if "n >= changeline", because
we want to include the scenario of author info at the first changed line
(which is the most typical scenario).
While on the subject, I see you've include the "if only ChangeLog
changes, ignore" heuristic. What about the case of the change in the
ChangeLog being a subtraction rather than an addition. I left that out
because all the scenarios I imagined that happening were a situation we
wanted to ignore any authorship change and just let the committer have
authorship (e.g., wholesale swap of ChangeLog.0 to ChangeLog.1, some
typo in the authorship line was corrected). I'm not sure your code
differentiates between the two (it looks to be searching just for
change), but really this is a very low likelihood of occurrence so maybe
it isn't worth toiling over that.
Dan
|
|
From: Eric S. R. <es...@th...> - 2017-10-21 11:42:35
|
Daniel J Sebald <dan...@ie...>: > I'd say no manual correction for that. The pool to search through is all > entries, 6000+. I was thinking the only manual adjustment might be the > substitution of "empty log message" as one of the last steps. That's only > 300 entries, much more reasonable to walk through. Agreed. You might want to take alook at the logic, though. I *think* my algorithm is equivalent to yours, but I could be wrong. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Eric S. R. <es...@th...> - 2017-10-21 11:39:14
|
"Bastian Märkisch" <bma...@we...>: > So it might be worth looking at the algorithm again. If the script > is available, I can try to help tweaking it. Otherwise, we would > have to come up with a list of manual corrections. In any case I > think we have a lot of work to do to verify the attributions. wget http://www.catb.org/~esr/gnuplot-conversion.tar.gz This will fetch you the conversion script and all the files needed to run it, into a directory named gnuplot-conversion. There's a README. The ChangeLog mining is in reposurgeon after around line 12417, in the function do_changelogs. I think it is well worth getting this right, as the code can then be re-used for other project conversions. Here is what it is intended to do: For each changeset in the repository conversion that includes a ChangeLog blob: 1. Fetch the ancestral version of ChangeLog, if it has one. Compare leading lines until you hit a pair that doesn't match. That line number beciomes the value of 'changeline'. If you find no ancestor, let changeline be 0. 2. Walk through the (newer) Changelog line file. Parse each attribution time for date and email address, soeing them in variables 'date' and 'addr'. 3. Break out of that loop when (a) the line you just parsed is at or after changeline, or (b) you reach EOF. There is a test in the test/ directory that shows this will correctly handle both the ordinary case of a new entry at the top of the file and the case where and entry is inserted after the top entry; see "Hilda J. Foonly" in liftlog.fi. The other files to look at are liftlog.tst (the reposurgeon script for this test) and liftlog.chk (the expected result). Note how in the .chk file the author slots of the last two commits have been set from the version of ChangLog in that commit. Run 'singletest liftlog' in test/ to rerun the test. (In case you've never seen one before, liftlog.fi and liflog.chk are git fast-export dumps of a small repository before and after surgery. These dumps capture the entire revision history and can be reconstituted into a live repo with fi-fast-import.) -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Daniel J S. <dan...@ie...> - 2017-10-21 08:13:12
|
On 10/21/2017 02:35 AM, "Bastian Märkisch" wrote:
>
>> Gesendet: Freitag, 20. Oktober 2017 um 23:01 Uhr
>> Von: "Eric S. Raymond" <es...@th...>
>> An: gnu...@li...
>> Betreff: News spin of repository conversion
>>
>> This one is 7a2323d15193540b226632f7b11ace3c524c63e3
>>
>> It was lifted using Daniel's enhanced algorithm for mining ChangeLog
>> files.
>
> Unfortunately there still seems to be something slightly odd about the algorithm:
> Ethan commited only a few changes by me in the beginning, yet I get:
> git log --committer=Ethan --author=Bastian --pretty="%H;%an;%cn;%cd;%s" | wc -l
> 132.
> I can also not remember ever having comitted a change by Ethan:
> git log --committer=Bastian --author=Ethan --pretty="%H;%an;%cn;%cd;%s" | wc -l
> 154
> I also suspect that the 53 commits by EAM authored by HBB are "false positives".
>
> Looking at some of these commits I see two cases where the algorithm fails:
> 1) A new ChangeLog entry without header is inserted. The algorithm seems to pick
> the wrong header from the ChangeLog.
> I thought that the backward search would identify these correctly.
Here's an example that Bastian is referring to. In gitg:
Ethan A M <xxxxxx@xxxxxx>
07/24/2017 12:00:02 PM +0000
Committed by: Bastian M <xxxxxx@xxxxxx>
07/27/2017 11:13:43 AM +0200
And the expanded diff hunk in the ChangeLog (the
first-change-search-back rule should hold):
sebald@ git-main $> git diff --unified=30
f0736a1c9285dc882528e55ed749ba6b2ac59ba0
e4e9a7594de9fc8a62f14d0dd4938f73aa781bb6
diff --git a/ChangeLog b/ChangeLog
index f46f85c..78f705d 100644
--- a/ChangeLog
+++ b/ChangeLog
@@ -1,35 +1,37 @@
2017-07-27 Bastian Maerkisch <bma...@we...>
* src/win/wd2d.cpp: Resize the swap chain buffers instead of
recreating the swap chain when the window size changes.
+ * src/win/wgraph.c term/win.trm: Default to Direct2D backend.
+
2017-07-24 Ethan A Merritt <merritt@u.washington.edu>
* src/datafile.c (df_open): Reject plot command if input and
output
both use the same data block. Prevents memory corruption /
segfault.
2017-07-24 Bastian Maerkisch <bma...@we...>
> 2) There have been additional corrections to / additions of previous ChangeLog entries
> in the same commit.
> Those are maybe hard to detect automatically.
Right. Or at least too much programming detail for little gain. I
suspect there aren't too many of these cases.
> So it might be worth looking at the algorithm again. If the script is available, I
> can try to help tweaking it. Otherwise, we would have to come up with a list of
> manual corrections. In any case I think we have a lot of work to do to verify the
> attributions.
I'd say no manual correction for that. The pool to search through is
all entries, 6000+. I was thinking the only manual adjustment might be
the substitution of "empty log message" as one of the last steps.
That's only 300 entries, much more reasonable to walk through.
Dan
>
> Bastian
>
> ------------------------------------------------------------------------------
> Check out the vibrant tech community on one of the world's most
> engaging tech sites, Slashdot.org! http://sdm.link/slashdot
> _______________________________________________
> gnuplot-beta mailing list
> gnu...@li...
> Membership management via: https://lists.sourceforge.net/lists/listinfo/gnuplot-beta
>
--
Dan Sebald
email: daniel(DOT)sebald(AT)ieee(DOT)org
URL: http://www(DOT)dansebald(DOT)com
|
|
From: Bastian M. <bma...@we...> - 2017-10-21 07:35:54
|
> Gesendet: Freitag, 20. Oktober 2017 um 23:01 Uhr > Von: "Eric S. Raymond" <es...@th...> > An: gnu...@li... > Betreff: News spin of repository conversion > > This one is 7a2323d15193540b226632f7b11ace3c524c63e3 > > It was lifted using Daniel's enhanced algorithm for mining ChangeLog > files. Unfortunately there still seems to be something slightly odd about the algorithm: Ethan commited only a few changes by me in the beginning, yet I get: git log --committer=Ethan --author=Bastian --pretty="%H;%an;%cn;%cd;%s" | wc -l 132. I can also not remember ever having comitted a change by Ethan: git log --committer=Bastian --author=Ethan --pretty="%H;%an;%cn;%cd;%s" | wc -l 154 I also suspect that the 53 commits by EAM authored by HBB are "false positives". Looking at some of these commits I see two cases where the algorithm fails: 1) A new ChangeLog entry without header is inserted. The algorithm seems to pick the wrong header from the ChangeLog. I thought that the backward search would identify these correctly. 2) There have been additional corrections to / additions of previous ChangeLog entries in the same commit. Those are maybe hard to detect automatically. So it might be worth looking at the algorithm again. If the script is available, I can try to help tweaking it. Otherwise, we would have to come up with a list of manual corrections. In any case I think we have a lot of work to do to verify the attributions. Bastian |
|
From: Achim G. <Str...@ne...> - 2017-10-20 19:33:09
|
Achim Gratz writes:
> So I finally did that today and there is still one tiny bit of
> regression left (I could have seen that in my original test cases
> already, but I didn't look closely enough). In the following example
> you'll see that the tics are only starting from 10^0=1 instead of from
> the smallest fractional power of of 10.
I don't really know what I'm doing, but this patch seems to fix this
issue. I don't know if it causes a regression somewhere else.
--8<---------------cut here---------------start------------->8---
--- origsrc/gnuplot-branch-5-2-stable/src/axis.c 2017-10-14 02:00:11.000000000 +0200
+++ src/gnuplot-branch-5-2-stable/src/axis.c 2017-10-20 21:23:34.189778800 +0200
@@ -1151,9 +1151,7 @@ gen_tics(struct axis *this, tic_callback
*/
step = eval_link_function(this->linked_to_primary, step);
end = eval_link_function(this->linked_to_primary, end);
- if (start <= 0)
- start = step;
- else
+ if (start > 0)
start = eval_link_function(this->linked_to_primary, start);
lmin = this->linked_to_primary->min;
lmax = this->linked_to_primary->max;
--8<---------------cut here---------------end--------------->8---
Regards,
Achim.
--
+<[Q+ Matrix-12 WAVE#46+305 Neuron microQkb Andromeda XTk Blofeld]>+
Samples for the Waldorf Blofeld:
http://Synth.Stromeko.net/Downloads.html#BlofeldSamplesExtra
|
|
From: Eric S. R. <es...@th...> - 2017-10-19 19:51:20
|
Bastian Märkisch <bma...@we...>: > > Can you be more specific about the time this happened, and in what way the > > git conversion fails to reflect the old history? > > > > If you look at > http://gnuplot.cvs.sourceforge.net/viewvc/gnuplot/gnuplot/?hideattic=0&pathr > ev=MAIN > there are a number of files marked "dead" deleted by Lars Hecking with a > commit > comment "Moved to ..." Similarily for /NeXT/, /beos/, /win/, /os2/. I was > wondering > if the history of these files could be prepended (again) to the history of > the destination > files. That is if that is easy enough. That part of the history was "lost" > some 18y ago ;) This probably didn't lose any history at all. If these were moved back to their original locations later, then both moves will be part of the gitspace history. Frankly, if this has gone wrong, I'd be very nervous about trying to fix it. CVS is flaky enough even when used as intended; moving around or otherwise messing with master files behind CVS's back tends to make very bad things happen. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Bastian M. <bma...@we...> - 2017-10-19 19:12:24
|
> > Another idea concerning the CVS conversion: at the early stages of > > the gnuplot CVS repo, files were moved from the top directory to > > sub-directories. That is they were deleted and added to CVS again. > > Hence the history before the move is kind of lost. While not super > > important, I wonder if this could be easily ammended now? > > Can you be more specific about the time this happened, and in what way the > git conversion fails to reflect the old history? > If you look at http://gnuplot.cvs.sourceforge.net/viewvc/gnuplot/gnuplot/?hideattic=0&pathr ev=MAIN there are a number of files marked "dead" deleted by Lars Hecking with a commit comment "Moved to ..." Similarily for /NeXT/, /beos/, /win/, /os2/. I was wondering if the history of these files could be prepended (again) to the history of the destination files. That is if that is easy enough. That part of the history was "lost" some 18y ago ;) Bastian > The way cvs-fast-export works *may* already solve this problem. It scans > masters in the attic directory - it has to, because that's where deleted files go. > In fact attic files are treated pretty much as though they were at their original, > pre-attic locations, so if a commit clique is found by matching time and > comment it does not matter that some of the file revisions are in the attic and > some are not. > -- > <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> > > My work is funded by the Internet Civil Engineering Institute: https://icei.org > Please visit their site and donate: the civilization you save might be your own. > > > > ---------------------------------------------------------------------------- -- > Check out the vibrant tech community on one of the world's most engaging > tech sites, Slashdot.org! http://sdm.link/slashdot > _______________________________________________ > gnuplot-beta mailing list > gnu...@li... > Membership management via: > https://lists.sourceforge.net/lists/listinfo/gnuplot-beta |
|
From: Hans-Bernhard B. <HBB...@t-...> - 2017-10-19 18:52:40
|
Am 19.10.2017 um 15:41 schrieb Petr Mikulik: > Slackware: > http://mirror.cslabs.clarkson.edu/slackware/slackware-3.4/source/xap/gnuplot/ > > 1993-09-25 00:23 626 008 gnuplot-3.5.tar.gz > > RedHat / Fedora: > http://pkgs.fedoraproject.org/repo/pkgs/gnuplot/ > 1999-11-07 16:57 1 319 233 gnuplot-3.7.1.tar.gz > 2002-02-25 20:40 1 399 872 gnuplot-3.7.2.tar.gz > 2002-12-12 14:00 1 418 889 gnuplot-3.7.3.tar.gz FWIW, the latter three (and all releases newer than these) are available right there at our own File download section on SourceForge, too :-) |
|
From: Achim G. <Str...@ne...> - 2017-10-19 18:26:14
|
> Did you mean to write > > set xtics 0.1 > > rather than 10? No, I really meant to write 10, which until version 5.2 meant that the tics would be spaced a decade (factor of 10) on a logscale axis. > For mxtics the number means frequency <freq> (i.e., > how many intervals to break into), while for xtics the number means > increment <incr> (i.e., the length of each interval). With xtics > increment of 10 there are no visible ticks on the x axis. There are, just not on the complete axis at the moment. Regards, Achim. -- +<[Q+ Matrix-12 WAVE#46+305 Neuron microQkb Andromeda XTk Blofeld]>+ Wavetables for the Waldorf Blofeld: http://Synth.Stromeko.net/Downloads.html#BlofeldUserWavetables |
|
From: Daniel J S. <dan...@ie...> - 2017-10-19 18:20:30
|
On 10/19/2017 01:05 PM, Achim Gratz wrote: > Achim Gratz writes: >> Well, once I see it appear in the repo I'll do a new build and run our >> plotting scripts @work against it. I had to rollback to 5.0.7 pretty >> quickly so I can't really say if there were no other problems, but >> nothing else that was glaringly obvious. > > So I finally did that today and there is still one tiny bit of > regression left (I could have seen that in my original test cases > already, but I didn't look closely enough). In the following example > you'll see that the tics are only starting from 10^0=1 instead of from > the smallest fractional power of of 10. > > --8<---------------cut here---------------start------------->8--- > set log x 10 > set mxtics 10 > set xtics 10 > set xrange [1e-3*pi:pi] > plot sin(x) > --8<---------------cut here---------------end--------------->8--- Did you mean to write set xtics 0.1 rather than 10? For mxtics the number means frequency <freq> (i.e., how many intervals to break into), while for xtics the number means increment <incr> (i.e., the length of each interval). With xtics increment of 10 there are no visible ticks on the x axis. Dan |
|
From: Achim G. <Str...@ne...> - 2017-10-19 18:06:01
|
Achim Gratz writes: > Well, once I see it appear in the repo I'll do a new build and run our > plotting scripts @work against it. I had to rollback to 5.0.7 pretty > quickly so I can't really say if there were no other problems, but > nothing else that was glaringly obvious. So I finally did that today and there is still one tiny bit of regression left (I could have seen that in my original test cases already, but I didn't look closely enough). In the following example you'll see that the tics are only starting from 10^0=1 instead of from the smallest fractional power of of 10. --8<---------------cut here---------------start------------->8--- set log x 10 set mxtics 10 set xtics 10 set xrange [1e-3*pi:pi] plot sin(x) --8<---------------cut here---------------end--------------->8--- Regards, Achim. -- +<[Q+ Matrix-12 WAVE#46+305 Neuron microQkb Andromeda XTk Blofeld]>+ Wavetables for the Waldorf Blofeld: http://Synth.Stromeko.net/Downloads.html#BlofeldUserWavetables |
|
From: sfeam <sf...@us...> - 2017-10-19 16:08:48
|
On Thursday, 19 October 2017 15:41:16 Petr Mikulik wrote: > I was looking into my old disc (copies) in order to find old gnuplot releases, > but I was not lucky with any pre-beta-3.4. Disc space was expensive that > time... > > Then searching web for old releases - unfortunately, most (all?) ftp sites > I remember for gnuplot source code download do not exist any longer. > Does somebody has them? > BTW, here is the ftp list: > https://groups.google.com/forum/#!topic/comp.graphics.gnuplot/rIrNBLNwdEI > showing > ftp://ftp.dartmouth.edu/pub/gnuplot/ > ftp://mirror.aarnet.edu.au/pub/gnuplot/ > ftp://ftp.irisa.fr/pub/gnuplot/ > ftp://ftp.ucc.ie/pub/gnuplot/ > ftp://ftp.gnuplot.vt.edu/pub/gnuplot/ > > http://members.theglobe.com/gnuplot/ > http://www.geocities.com/SiliconValley/Foothills/6647/ > http://mirror.aarnet.edu.au/pub/gnuplot/ > > > Finally, I have found old gnuplot source codes at two Linux distributions we > were using long time ago - Slackware and RedHat (diskette-based distros :-). > > Slackware: > http://mirror.cslabs.clarkson.edu/slackware/slackware-3.4/source/xap/gnuplot/ > 1993-09-25 00:23 626 008 gnuplot-3.5.tar.gz > > RedHat / Fedora: > http://pkgs.fedoraproject.org/repo/pkgs/gnuplot/ > 1999-11-07 16:57 1 319 233 gnuplot-3.7.1.tar.gz > 2002-02-25 20:40 1 399 872 gnuplot-3.7.2.tar.gz > 2002-12-12 14:00 1 418 889 gnuplot-3.7.3.tar.gz Good work tracking those down. It's a very minor point, but the internal timestamps in version.c for those releases are gnuplot-3.7.1 "Fri Oct 22 18:00:00 BST 1999" gnuplot-3.7.2 "Sat Jan 19 15:23:37 GMT 2002" gnuplot-3.7.3 "Thu Dec 12 13:00:00 GMT 2002" Ah nostalgia. I think 3.7.1 was the first version I used. Ethan > > I have adjusted the dates according to the oldest file in the package, > and put them here: > https://www.physics.muni.cz/~mikulik/gnuplot/ > > Can you add those 4 .tar.gz files to the git repo? > > --- > Petr Mikulik |
|
From: Petr M. <mi...@ph...> - 2017-10-19 13:41:29
|
I was looking into my old disc (copies) in order to find old gnuplot releases, but I was not lucky with any pre-beta-3.4. Disc space was expensive that time... Then searching web for old releases - unfortunately, most (all?) ftp sites I remember for gnuplot source code download do not exist any longer. Does somebody has them? BTW, here is the ftp list: https://groups.google.com/forum/#!topic/comp.graphics.gnuplot/rIrNBLNwdEI showing ftp://ftp.dartmouth.edu/pub/gnuplot/ ftp://mirror.aarnet.edu.au/pub/gnuplot/ ftp://ftp.irisa.fr/pub/gnuplot/ ftp://ftp.ucc.ie/pub/gnuplot/ ftp://ftp.gnuplot.vt.edu/pub/gnuplot/ http://members.theglobe.com/gnuplot/ http://www.geocities.com/SiliconValley/Foothills/6647/ http://mirror.aarnet.edu.au/pub/gnuplot/ Finally, I have found old gnuplot source codes at two Linux distributions we were using long time ago - Slackware and RedHat (diskette-based distros :-). Slackware: http://mirror.cslabs.clarkson.edu/slackware/slackware-3.4/source/xap/gnuplot/ 1993-09-25 00:23 626 008 gnuplot-3.5.tar.gz RedHat / Fedora: http://pkgs.fedoraproject.org/repo/pkgs/gnuplot/ 1999-11-07 16:57 1 319 233 gnuplot-3.7.1.tar.gz 2002-02-25 20:40 1 399 872 gnuplot-3.7.2.tar.gz 2002-12-12 14:00 1 418 889 gnuplot-3.7.3.tar.gz I have adjusted the dates according to the oldest file in the package, and put them here: https://www.physics.muni.cz/~mikulik/gnuplot/ Can you add those 4 .tar.gz files to the git repo? --- Petr Mikulik |
|
From: Eric S. R. <es...@th...> - 2017-10-19 01:16:33
|
Hans-Bernhard Bröker <HBB...@t-...>: > Am 19.10.2017 um 01:30 schrieb Eric S. Raymond: > > >Next question is, shall I nuke the tags that don't refer to a gitspace > >revision? Checking these out probably won't reflect what was intended, due > >to missing CVS tags in some masters that should have them. > > Most of those can go, but some, IMHO, do deserve being kept, if only on a > 'best-effort' basis. > > Basically, all tags with a date stamp in their name really should go. > > > GNUPLOT_RELEASE_3_7_0 > > GNUPLOT_RELEASE_4_0_0 > > GNUPLOT_RELEASE_4_0_1 > > GNUPLOT_RELEASE_4_0_2 > > Release_4_6_3 > > These, OTOH, really have to be kept. If they're no longer viable in the CVS > repository, that is a problem worthy of attention all by itself. Could you > elaborate what's wrong about them? OK. Release_4_6_3 was included on that list by mistake; Ethan pointed that out and in the next spin it will be renamed tp 4.6.3 The others are victims of an all-too-common form of metadata damage in CVS repositories. The way tagging is *supposed* to work is than when you create a tag it gets added to every master. When this happens, cvs-fast-export notices that the tag is complete. Complete tags get associated with gitspace changesets. However, for various unclear reasons that can be grouped under "CVS is a rickety pile of kludges that should have been strangled in its crib", CVS sometimes fails to create tags in every master where it should. When this happens - the tag is incomplete - checking out at that tag can omit content that was intended to be included when the tag was made. Thete's no way to know if this will occur. That applies to these tags: GNUPLOT_RELEASE_3_7_0 GNUPLOT_RELEASE_4_0_0 GNUPLOT_RELEASE_4_0_1 GNUPLOT_RELEASE_4_0_2 Because incomplete tags are not reliable, cvs-fast-export does not try to associate them to a gitspace changeset. instead, it creates a synthetic commit for each one with a comment explaining that it is incomplete. I think the best thing to do would be to throw these out and reconstruct correct gitspace tags for these versions by looking at where src/versions.c changes. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Hans-Bernhard B. <HBB...@t-...> - 2017-10-19 00:12:50
|
Am 19.10.2017 um 01:30 schrieb Eric S. Raymond: > Next question is, shall I nuke the tags that don't refer to a gitspace > revision? Checking these out probably won't reflect what was intended, due > to missing CVS tags in some masters that should have them. Most of those can go, but some, IMHO, do deserve being kept, if only on a 'best-effort' basis. Basically, all tags with a date stamp in their name really should go. > GNUPLOT_RELEASE_3_7_0 > GNUPLOT_RELEASE_4_0_0 > GNUPLOT_RELEASE_4_0_1 > GNUPLOT_RELEASE_4_0_2 > Release_4_6_3 These, OTOH, really have to be kept. If they're no longer viable in the CVS repository, that is a problem worthy of attention all by itself. Could you elaborate what's wrong about them? |