You can subscribe to this list here.
| 2001 |
Jan
|
Feb
(1) |
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
|
Sep
|
Oct
|
Nov
|
Dec
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2002 |
Jan
(1) |
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
|
Nov
(1) |
Dec
|
| 2003 |
Jan
|
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
(1) |
Aug
(1) |
Sep
|
Oct
(83) |
Nov
(57) |
Dec
(111) |
| 2004 |
Jan
(38) |
Feb
(121) |
Mar
(107) |
Apr
(241) |
May
(102) |
Jun
(190) |
Jul
(239) |
Aug
(158) |
Sep
(184) |
Oct
(193) |
Nov
(47) |
Dec
(68) |
| 2005 |
Jan
(190) |
Feb
(105) |
Mar
(99) |
Apr
(65) |
May
(92) |
Jun
(250) |
Jul
(197) |
Aug
(128) |
Sep
(101) |
Oct
(183) |
Nov
(186) |
Dec
(42) |
| 2006 |
Jan
(102) |
Feb
(122) |
Mar
(154) |
Apr
(196) |
May
(181) |
Jun
(281) |
Jul
(310) |
Aug
(198) |
Sep
(145) |
Oct
(188) |
Nov
(134) |
Dec
(90) |
| 2007 |
Jan
(134) |
Feb
(181) |
Mar
(157) |
Apr
(57) |
May
(81) |
Jun
(204) |
Jul
(60) |
Aug
(37) |
Sep
(17) |
Oct
(90) |
Nov
(122) |
Dec
(72) |
| 2008 |
Jan
(130) |
Feb
(108) |
Mar
(160) |
Apr
(38) |
May
(83) |
Jun
(42) |
Jul
(75) |
Aug
(16) |
Sep
(71) |
Oct
(57) |
Nov
(59) |
Dec
(152) |
| 2009 |
Jan
(73) |
Feb
(213) |
Mar
(67) |
Apr
(40) |
May
(46) |
Jun
(82) |
Jul
(73) |
Aug
(57) |
Sep
(108) |
Oct
(36) |
Nov
(153) |
Dec
(77) |
| 2010 |
Jan
(42) |
Feb
(171) |
Mar
(150) |
Apr
(6) |
May
(22) |
Jun
(34) |
Jul
(31) |
Aug
(38) |
Sep
(32) |
Oct
(59) |
Nov
(13) |
Dec
(62) |
| 2011 |
Jan
(114) |
Feb
(139) |
Mar
(126) |
Apr
(51) |
May
(53) |
Jun
(29) |
Jul
(41) |
Aug
(29) |
Sep
(35) |
Oct
(87) |
Nov
(42) |
Dec
(20) |
| 2012 |
Jan
(111) |
Feb
(66) |
Mar
(35) |
Apr
(59) |
May
(71) |
Jun
(32) |
Jul
(11) |
Aug
(48) |
Sep
(60) |
Oct
(87) |
Nov
(16) |
Dec
(38) |
| 2013 |
Jan
(5) |
Feb
(19) |
Mar
(41) |
Apr
(47) |
May
(14) |
Jun
(32) |
Jul
(18) |
Aug
(68) |
Sep
(9) |
Oct
(42) |
Nov
(12) |
Dec
(10) |
| 2014 |
Jan
(14) |
Feb
(139) |
Mar
(137) |
Apr
(66) |
May
(72) |
Jun
(142) |
Jul
(70) |
Aug
(31) |
Sep
(39) |
Oct
(98) |
Nov
(133) |
Dec
(44) |
| 2015 |
Jan
(70) |
Feb
(27) |
Mar
(36) |
Apr
(11) |
May
(15) |
Jun
(70) |
Jul
(30) |
Aug
(63) |
Sep
(18) |
Oct
(15) |
Nov
(42) |
Dec
(29) |
| 2016 |
Jan
(37) |
Feb
(48) |
Mar
(59) |
Apr
(28) |
May
(30) |
Jun
(43) |
Jul
(47) |
Aug
(14) |
Sep
(21) |
Oct
(26) |
Nov
(10) |
Dec
(2) |
| 2017 |
Jan
(26) |
Feb
(27) |
Mar
(44) |
Apr
(11) |
May
(32) |
Jun
(28) |
Jul
(75) |
Aug
(45) |
Sep
(35) |
Oct
(285) |
Nov
(99) |
Dec
(16) |
| 2018 |
Jan
(8) |
Feb
(8) |
Mar
(42) |
Apr
(35) |
May
(23) |
Jun
(12) |
Jul
(16) |
Aug
(11) |
Sep
(8) |
Oct
(16) |
Nov
(5) |
Dec
(8) |
| 2019 |
Jan
(9) |
Feb
(28) |
Mar
(4) |
Apr
(10) |
May
(7) |
Jun
(4) |
Jul
(4) |
Aug
|
Sep
(4) |
Oct
|
Nov
(23) |
Dec
(3) |
| 2020 |
Jan
(19) |
Feb
(3) |
Mar
(22) |
Apr
(17) |
May
(10) |
Jun
(69) |
Jul
(18) |
Aug
(23) |
Sep
(25) |
Oct
(11) |
Nov
(20) |
Dec
(9) |
| 2021 |
Jan
(1) |
Feb
(7) |
Mar
(9) |
Apr
|
May
(1) |
Jun
(8) |
Jul
(6) |
Aug
(8) |
Sep
(7) |
Oct
|
Nov
(2) |
Dec
(23) |
| 2022 |
Jan
(23) |
Feb
(9) |
Mar
(9) |
Apr
|
May
(8) |
Jun
(1) |
Jul
(6) |
Aug
(8) |
Sep
(30) |
Oct
(5) |
Nov
(4) |
Dec
(6) |
| 2023 |
Jan
(2) |
Feb
(5) |
Mar
(7) |
Apr
(3) |
May
(8) |
Jun
(45) |
Jul
(8) |
Aug
|
Sep
(2) |
Oct
(14) |
Nov
(7) |
Dec
(2) |
| 2024 |
Jan
(4) |
Feb
(4) |
Mar
|
Apr
(7) |
May
(2) |
Jun
(1) |
Jul
|
Aug
(5) |
Sep
|
Oct
|
Nov
(4) |
Dec
(14) |
| 2025 |
Jan
(22) |
Feb
(6) |
Mar
(5) |
Apr
(14) |
May
(6) |
Jun
(11) |
Jul
(19) |
Aug
|
Sep
(17) |
Oct
(1) |
Nov
(2) |
Dec
(18) |
| 2026 |
Jan
|
Feb
|
Mar
(5) |
Apr
|
May
(2) |
Jun
(1) |
Jul
(6) |
Aug
(1) |
Sep
|
Oct
|
Nov
|
Dec
|
|
From: sfeam <sf...@us...> - 2017-10-14 20:04:10
|
On Saturday, 14 October 2017 14:15:50 Eric S. Raymond wrote: > I'm making good progress on the git conversion procedure. I expect to > have it wrapped up and ready to go tonight or tomorrow. All I will > need to pull the trigger at that point is > > (a) Daniel Sebald's authorship map. > > (b) Push privileges on the git hosting site. I was thinking that you would create and populate the git repository somewhere convenient to you, where you have control over everything up to and including blowing it away and starting over. Then when it seems to be successfully set up we would clone it onto SourceForge. I can add you to the project "admin" or "developer" groups if you have a SourceForge user ID. More fine-grained management of privileges on the site is something I have never found proper documentation for. I created a placeholder by hitting a button "Add new git" on the admin tab https://sourceforge.net/p/gnuplot/git-main/ref/master/ but I don't know exactly which subset of users associated with the project can push to it. I am guessing it's everyone in the "developer" group. Please bear with me if I am misunderstanding what needs to be done. I have used git only to clone and use an existing repository; I've never set one up. > Ethan, you sent me three tarballs of archaic releases: > > gnuplot-1.10A.tar.gz > gnuplot-2.0.tar.gz > gnuplot-3.5.tar.gz > > I have the machinery set up to graft these onto the front of the git > repository, but I nees one piece of metadata for each: the release > date. The timestamps in the tarballs seem to have been clobbered > at some point. >From the self-reported date in the source file version.c gnuplot 1.1.0A "Thu May 18 21:57:24 MST 1989" gnuplot 2.0.0 "Wed Mar 7 22:18:59 EST 1990" gnuplot 3.5 "Fri Aug 27 05:21:33 GMT 1993" Yeah the timestamps probably don't mean much. I reconstructed the source trees from multipart shar archives originally posed to usenet group comp.sources.misc and stored in that form in the gnuplot-historical subdirectory on sf.net. I see there are also shar files there for "gnuplot-3.0", so I may try to unpack and put those in a tarball also. Ethan Ethan |
|
From: <es...@th...> - 2017-10-14 18:15:58
|
I'm making good progress on the git conversion procedure. I expect to have it wrapped up and ready to go tonight or tomorrow. All I will need to pull the trigger at that point is (a) Daniel Sebald's authorship map. (b) Push privileges on the git hosting site. Ethan, you sent me three tarballs of archaic releases: gnuplot-1.10A.tar.gz gnuplot-2.0.tar.gz gnuplot-3.5.tar.gz I have the machinery set up to graft these onto the front of the git repository, but I nees one piece of metadata for each: the release date. The timestamps in the tarballs seem to have been clobbered at some point. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> Don't think of it as `gun control', think of it as `victim disarmament'. If we make enough laws, we can all be criminals. |
|
From: Daniel J S. <dan...@ie...> - 2017-10-13 23:07:53
|
On 10/13/2017 05:20 PM, Ethan A Merritt wrote: > On Friday, 13 October, 2017 16:44:59 Daniel J Sebald wrote: > > > Ethan mentioned offline old CVS branches, that brings to mind something > > about git worth mentioning for those maintainers not familiar with git. > > git uses terminology "branch" quite often, but after working with git > > for a while one will realize the "branch" analogy fails a bit after > > merges are done. More accurately, git pushes the heads along. It's > > important noting this because once two branch heads are merged, the > > names once associated with those branches are lost and one doesn't know > > what branch was what at the time they were under development. > > So post-conversion, what is the git terminology that distinguishes > the stable 5.2 branch from the main 5.3 branch? > If I want to apply a patch specifically to the 5.2 "branch" > (clone something / modify it / merge back) > what are the git commands I use to do that so that it doesn't > affect the main development "branch"? > > Ethan You may create branches anywhere in the repository. The first step would be to checkout the 5.2 branch with git checkout branch-5-2-stable with the requirement that there be no modified files or staged modified files. (Plain "git checkout FILENAME" will discard file changes and return the file to the repository.) Now once in branch-5-2-stable, do a branch git checkout -b b52_patch and any modifications and changesets one makes at that point end up on "branch"/"head" b52_patch. I've done the above in the test repository, and inquiring the branch info tells me: sebald@ ~/gnuplot/test_repository/gnuplot $ git branch * b52_patch branch-5-2-stable master I think what the above is telling us is b52_patch is the current active branch, branch-5-2-stable was the prior branch, and prior to that was master. If one does a "merge" from within b52_patch, I believe git assumes the prior branch "branch-5-2-stable" is to be the merge point. However, one can merge anywhere, I think, just by specifying an object (tag, head name, probably SHA number or however many first X unique numbers). Dan |
|
From: Ethan A M. <merritt@u.washington.edu> - 2017-10-13 22:37:13
|
On Friday, 13 October, 2017 16:44:59 Daniel J Sebald wrote: > Ethan mentioned offline old CVS branches, that brings to mind something > about git worth mentioning for those maintainers not familiar with git. > git uses terminology "branch" quite often, but after working with git > for a while one will realize the "branch" analogy fails a bit after > merges are done. More accurately, git pushes the heads along. It's > important noting this because once two branch heads are merged, the > names once associated with those branches are lost and one doesn't know > what branch was what at the time they were under development. So post-conversion, what is the git terminology that distinguishes the stable 5.2 branch from the main 5.3 branch? If I want to apply a patch specifically to the 5.2 "branch" (clone something / modify it / merge back) what are the git commands I use to do that so that it doesn't affect the main development "branch"? Ethan > > For example, say I create a branch from "master" called "sebald001" with > > git checkout -b sebald001 > > I then make some commits on "sebald001", while on "master" some other > commits are made by others. If the two are merged, it's not obvious > what changes were done at the time on "sebald001" versus those that were > done at the time on "master". > > The reason this is important is because often git projects have a rule > of no modifications on the "master" branch, only merges. So, even the > smallest of mods are done as a branch and then merged. Whether you want > to do that here, where there are a lot of people submitting changesets > and patch sets is worth discussion. I believe that any submitted > changesets can be rebased locally by the maintainers to fit any model > you decide upon and then pushed to the main repository. > > In any case, there are some logistical and procedural things the > maintainers will have to decide on. I just want to point out that > although git gives a ton of flexibility in terms of merging branches it > can be real confusing after the fact to know exactly how things > progressed when there are intertwined multiple merged branches. > > Dan > > ------------------------------------------------------------------------------ > Check out the vibrant tech community on one of the world's most > engaging tech sites, Slashdot.org! http://sdm.link/slashdot > _______________________________________________ > gnuplot-beta mailing list > gnu...@li... > Membership management via: https://lists.sourceforge.net/lists/listinfo/gnuplot-beta -- Ethan A Merritt Biomolecular Structure Center, K-428 Health Sciences Bldg MS 357742, University of Washington, Seattle 98195-7742 |
|
From: Daniel J S. <dan...@ie...> - 2017-10-13 22:01:55
|
Ethan mentioned offline old CVS branches, that brings to mind something about git worth mentioning for those maintainers not familiar with git. git uses terminology "branch" quite often, but after working with git for a while one will realize the "branch" analogy fails a bit after merges are done. More accurately, git pushes the heads along. It's important noting this because once two branch heads are merged, the names once associated with those branches are lost and one doesn't know what branch was what at the time they were under development. For example, say I create a branch from "master" called "sebald001" with git checkout -b sebald001 I then make some commits on "sebald001", while on "master" some other commits are made by others. If the two are merged, it's not obvious what changes were done at the time on "sebald001" versus those that were done at the time on "master". The reason this is important is because often git projects have a rule of no modifications on the "master" branch, only merges. So, even the smallest of mods are done as a branch and then merged. Whether you want to do that here, where there are a lot of people submitting changesets and patch sets is worth discussion. I believe that any submitted changesets can be rebased locally by the maintainers to fit any model you decide upon and then pushed to the main repository. In any case, there are some logistical and procedural things the maintainers will have to decide on. I just want to point out that although git gives a ton of flexibility in terms of merging branches it can be real confusing after the fact to know exactly how things progressed when there are intertwined multiple merged branches. Dan |
|
From: Daniel J S. <dan...@ie...> - 2017-10-13 20:14:57
|
On 10/13/2017 01:48 PM, Ethan A Merritt via gnuplot-beta wrote: > On Friday, 13 October, 2017 08:56:36 Eric S. Raymond wrote: > > By the way, while working on this I found a bad date 2206-07-27 > > in term/lua/ChangeLog. That probably wants to be 2006; somebody > > should fix it. > > There you open a can of worms I'd just as soon leave closed. > That file no longer exists in the set of files obtained by generic > cvs checkout. Yes it can be found in the repository as a "dead file", > but I don't think it's worth my time to go mucking about in the > graveyard of dead files to correct typos. > > Ethan What files no longer exist? The 2206 currently appears in the file ChangeLog.1 of the current CVS repository. If you were to change that file, I don't think it would have any ramifications on the git conversion. That date info wouldn't be used for anything. The 2206 still appears in the historical repository record; no changing that. But again, those dates aren't used for anything in the conversion scheme. I would say that making any changes to ChangeLog files--typos, dates, etc.--is better done before the git conversion than after. Do we really want to be doing such changes under the new git repository? I can imagine regrouping the ChangeLog files, even deleting the ChangeLog files if later you think git log -p ChangeLog | less is sufficient for searching old ChangeLog mods. (Can always retrieve old ChangeLog files for local use.) But to be cleaning up ChangeLog files that are sort of deprecated under the new source control seems odd. It might be better freezing them at that point. Dan |
|
From: Eric S. R. <es...@th...> - 2017-10-13 19:11:51
|
Ethan A Merritt <sf...@us...>: > On Friday, 13 October, 2017 08:56:36 Eric S. Raymond wrote: > > By the way, while working on this I found a bad date 2206-07-27 > > in term/lua/ChangeLog. That probably wants to be 2006; somebody > > should fix it. > > There you open a can of worms I'd just as soon leave closed. > That file no longer exists in the set of files obtained by generic > cvs checkout. Yes it can be found in the repository as a "dead file", > but I don't think it's worth my time to go mucking about in the > graveyard of dead files to correct typos. OK. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Eric S. R. <es...@th...> - 2017-10-13 19:11:16
|
Daniel J Sebald <dan...@ie...>: > Oh, OK. Yeah, I've found a command that does it without even having to > change the code (posted at the patch-tracker): > > git log --format='commit <%cE!%cIZ>' -p --unified=50 ChangeLog > > ChangeLog.diff > git_changelog_author ChangeLog.diff > author.txt > > The "commit" in the first column is what the utility is searching for. So > the list then looks like this: > > <sfeam!2017-09-04T05:57:37+00:00Z> 2017-09-03 Daniel J Sebxxx > <xxxxxx@xxxxxx> > <sfeam!2017-09-03T23:28:50+00:00Z> 2017-09-03 Ethan A Merrxxx > <xxxxxx@xxxxxx> > <sfeam!2017-09-03T03:35:30+00:00Z> 2017-09-02 Ethan A Merrxxx > <xxxxxx@xxxxxx> > <markisch!2017-09-01T05:21:00+00:00Z> 2017-09-01 Bastian Maerkixxx > <xxxxxx@xxxxxx> > <sfeam!2017-08-31T23:03:55+00:00Z> 2017-08-31 Martin Satuxxx <xxxxxx@xxxxxx> > <sfeam!2017-08-30T20:30:12+00:00Z> 2017-08-24 Ethan A Merrxxx > <xxxxxx@xxxxxx> > > which has the person who committed the changeset and author. Nice. OK, I can easily massage that report into a set of reposurgeon commands that will patch in the author fields. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Eric S. R. <es...@th...> - 2017-10-13 19:06:36
|
Ethan A Merritt <sf...@us...>: > I can believe that 12 hours from the nominal date stamp misses many > of my commits, given that if I label an entry with today's date > anything I do in the afternoon probably gets a next-day timestamp > because of the time-zone difference. > > Even so that seems like too low a hit rate on the match-up. > My experience when I've tried to match a reported bug to a likely > culpable patch by taking the date in the ChangeLog and manually > searching for a matching repository commit date in "cvs view" > there is a high chance of finding it easily (same-day or next-day > match). Certainly I've not seen a 99% failure rate. > > Wait, hang on - what do you mean by "contributor ID"? In this entry 2009-03-12 Ethan A Merritt <merritt@u.washington.edu> * src/fit.c (fit_command): Replace bogus initialization of dummy_token[] with explicit declaration. Bug #2657599 the contributor-ID is everything on the header line after the date stamp. > I think you can only match the filename and the commit time And that's the code I wrote does. The committer-ID isn't used to *find* a match between log entry and commit, it's used to fill the author field of the commit *if an entry match on timestamp and pathset is found*. > Another thought... What do you mean by "match the path set"? > CVS commits occur one-by-one, so there isn't really a "set" to search for. > I can easily believe that the set of filenames listed in the ChangeLog > omits some that were committed by at the same time by mistake or > intent (trivial change to comment or some such). I think it makes sense > only to search for a match to one file at a time, not a set of files. Under that assumption the entire procedure is doomed. The heavy lifting in a CVS-to-git conversion is precisely the part where you recognize CVS single-file commits and group them into git changesets. You do this by recognizing identical committer metadata and comments and timestamps that are within a certain window of time difference, usually 15 minutes. By the time the Changelog-mining code sees the repository those changesets have already been grouped. What you want to do is set authorship for the *changesets*, not the compoment per-file commits that don't actually exist any more. What I've found out is that the ChangeLogs themselves are not sufficient for that. Daniel Sebald is working on a different approach based on examining diffs that he thinks might have better results. > Could there be an issue with mis-match of branch identifiers? > I.e. a patch is often applied to the main (development) branch first, > and then applied much later to the current release branch. > The ChangeLog entries and commit messages are often identical > but for the time stamp. Could your automated search be confused by > finding a potential match in one branch but then comparing it to > a timestamp from a different branch? It looks for every potential match that is close enough in time, so one of those would probably get filled in but not the other. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Ethan A M. <sf...@us...> - 2017-10-13 18:50:56
|
On Friday, 13 October, 2017 08:56:36 Eric S. Raymond wrote: > By the way, while working on this I found a bad date 2206-07-27 > in term/lua/ChangeLog. That probably wants to be 2006; somebody > should fix it. There you open a can of worms I'd just as soon leave closed. That file no longer exists in the set of files obtained by generic cvs checkout. Yes it can be found in the repository as a "dead file", but I don't think it's worth my time to go mucking about in the graveyard of dead files to correct typos. Ethan |
|
From: Daniel J S. <dan...@ie...> - 2017-10-13 18:44:19
|
On 10/13/2017 01:21 PM, Eric S. Raymond wrote: > Daniel J Sebald <dan...@ie...>: >> On 10/13/2017 07:56 AM, Eric S. Raymond wrote: >>> You haven't heard from me in a couple of days because I decided to >>> build a ChangeLog parser into reposurgeon and have been working on >>> that. >>> >>> It's not an easy problem, and I have doubts that the information >>> extracted that way will either cover a lot of commits or be of >>> useful quality. >> >> I have done such a thing with the post git-converted repository already. >> Code is here: >> >> https://sourceforge.net/p/gnuplot/patches/763/ >> >> I believe it is rather accurate, more accurate than the method you are >> describing which attempts to pull info from the ChangeLog entries alone. >> >> The methodology in the link above has the added, important detail >> (particular to the gnuplot project) of working from the diff hunks of the >> ChangeLog file itself for every git-translated changeset. It hinges on the >> fact that Ethan, et al. have consistently modified the proper location in >> the ChangeLog file for every CVS modification they've made. It's that action >> over the years which makes it accurate, i.e., the method traces the >> incremental changes made by the maintainers. >> >> Parsing the ChangeLog on its own won't work very well because, as I pointed >> out, the entries have been generated piecemeal. Ethan, et al. would often >> make a mod that goes back several days or possibly weeks pertaining to some >> checkin attributed to various contributors. Chronological order is not >> maintained; the mod maintainers make can be weeks apart from the date listed >> in the ChangeLog entry. >> >> That's why relying on the particular mod the maintainers made is better. >> >> The utility generates a list similar to the following: >> >> 87ca52fd56cf60be023642ad2664390affd5d896 2017-09-03 Daniel J Sebxxx >> <xxxxxx@xxxxxx> >> 778fe3d377116dc6d8c4f855372795bea5339e62 2017-09-03 Ethan A Merrxxx >> <xxxxxx@xxxxxx> >> c67c414e1604b20c562451ca6a2d6e9dca71561f 2017-09-02 Ethan A Merrxxx >> <xxxxxx@xxxxxx> >> 357f0cf68d691f3f76aade64053a6791f2e28686 2017-09-01 Bastian Maerkixxx >> <xxxxxx@xxxxxx> >> 984f9ad6c62f1c3ade9a41b46b5e034a2cb01c0a 2017-08-31 Martin Satuxxx >> <xxxxxx@xxxxxx> >> a0b903cbcae615cef9c0c0e3942c78ab150e3c82 2017-08-24 Ethan A Merrxxx >> <xxxxxx@xxxxxx> >> >> I pointed out in a previous email that the date in the above list is of no >> value for the git repository. There will not be good correlation between >> the date above and the actual date of the git changeset associated with >> 87ca52fd56cf60be023642ad2664390affd5d896, for example. All that needs to be >> done is some automated git command that will associate the author >> information in the above list with the gitID changeset in the list. > > I see your point. Unfortunately, I can't integrate this technique into > my conversion pipeline as-is, because > > (a) it relies, as you note, on a regularity in metadata specific to your > project, and > > (b) My tools know nothing about SHA-1 IDs! Those are specific to an > instantiated repository's hash chains. Yes, it's particular to the repository; no transferring of lists across platforms, conversions, etc. > However, there may be a way around this. > > Can you generate your report with the hashes replaced by action stamps? > > An action-stamp is a kind of commit ID my tools understand. It has > this form: > > <committer-id!yyyy-mm-ddThh:mm:ssZ> > > That is, it consists of a committer ID email address, followed by an > RFC3339-format representation of the commit date. > > These are not necessarily unique per repository, but they usually are. > > This command will generate an action stamp from a SHA-1 ID: > > git log --format='<%cE!%cIZ>' -1 $1 Oh, OK. Yeah, I've found a command that does it without even having to change the code (posted at the patch-tracker): git log --format='commit <%cE!%cIZ>' -p --unified=50 ChangeLog > ChangeLog.diff git_changelog_author ChangeLog.diff > author.txt The "commit" in the first column is what the utility is searching for. So the list then looks like this: <sfeam!2017-09-04T05:57:37+00:00Z> 2017-09-03 Daniel J Sebxxx <xxxxxx@xxxxxx> <sfeam!2017-09-03T23:28:50+00:00Z> 2017-09-03 Ethan A Merrxxx <xxxxxx@xxxxxx> <sfeam!2017-09-03T03:35:30+00:00Z> 2017-09-02 Ethan A Merrxxx <xxxxxx@xxxxxx> <markisch!2017-09-01T05:21:00+00:00Z> 2017-09-01 Bastian Maerkixxx <xxxxxx@xxxxxx> <sfeam!2017-08-31T23:03:55+00:00Z> 2017-08-31 Martin Satuxxx <xxxxxx@xxxxxx> <sfeam!2017-08-30T20:30:12+00:00Z> 2017-08-24 Ethan A Merrxxx <xxxxxx@xxxxxx> which has the person who committed the changeset and author. Nice. Dan |
|
From: Ethan A M. <sf...@us...> - 2017-10-13 18:32:24
|
On Friday, 13 October, 2017 12:40:48 Eric S. Raymond wrote: > Alas, automated ChangeLog mining to fill in author slots doesn't work > well enough to be usable. > > The algorithm I ended up implementing digests the ChangeLogs into a > set of tuples each consiring of a date stamp (with no time part), a > contributor ID, and a set of paths. It then walks through the entry > list, looking for commits that match the path set. Then it filters > for commits close in time to the entry timestamp. > > "Close" is by default 12 hours to either side of the commit stamp. I can believe that 12 hours from the nominal date stamp misses many of my commits, given that if I label an entry with today's date anything I do in the afternoon probably gets a next-day timestamp because of the time-zone difference. Even so that seems like too low a hit rate on the match-up. My experience when I've tried to match a reported bug to a likely culpable patch by taking the date in the ChangeLog and manually searching for a matching repository commit date in "cvs view" there is a high chance of finding it easily (same-day or next-day match). Certainly I've not seen a 99% failure rate. Wait, hang on - what do you mean by "contributor ID"? Many of the patches I commit are from other contributors. My name does not appear in the ChangeLog entry. The CVS repository presumably lists me as the source of the commit and offers no hint that the ChangeLog entry shows someone else as the contributor. I think you can only match the filename and the commit time (and with some human intelligence the correlation between the description in the ChangeLog and the short text in the commit message). As a sanity check or additional guide it might be reasonable to use a mapping of contributor name/email -> likely committer. Another thought... What do you mean by "match the path set"? CVS commits occur one-by-one, so there isn't really a "set" to search for. I can easily believe that the set of filenames listed in the ChangeLog omits some that were committed by at the same time by mistake or intent (trivial change to comment or some such). I think it makes sense only to search for a match to one file at a time, not a set of files. Could there be an issue with mis-match of branch identifiers? I.e. a patch is often applied to the main (development) branch first, and then applied much later to the current release branch. The ChangeLog entries and commit messages are often identical but for the time stamp. Could your automated search be confused by finding a potential match in one branch but then comparing it to a timestamp from a different branch? Ethan > The > window can be extended on the future side to allow for delay in > merging patches. > > $ reposurgeon "read ." "changelogs" > reposurgeon: Fills 123 of 49627 authorship slots from 4915 ChangeLog entries > $ reposurgeon "read ." "changelogs 36" > reposurgeon: Fills 468 of 49627 authorship slots from 4915 ChangeLog entries > $ reposurgeon "read ." "changelogs 72" > reposurgeon: Fills 498 of 49627 authorship slots from 4915 ChangeLog entries > $ reposurgeon "read ." "changelogs 120" > reposurgeon: Fills 515 of 49627 authorship slots from 4915 ChangeLog entries > > The argument of changelogs is a count of hours to extend the > closeness window by. As expected, increasing the window yields > more matches. > > However, even with a very long window the match rate never goes above > 1% of commits. That is noise level. > > I think the reason is hinted at by the large disparity (about 10:1) > between commit cliques and ChangeLog entries. What this tells us is > that a typical ChangeLog entry corresponds not to one commit clique > but to several. There's no algorithmic way to know what the boundaries > are. > > If such annotations are going to be made they will need a human eye > and hand comparing at each of 4916 ChangeLog entries against the > commit history. > > That could be done with reposurgeon; I ran some numbers and it's > probably about 80 hours of hand-work. No thanks... > |
|
From: Eric S. R. <es...@th...> - 2017-10-13 18:22:02
|
Daniel J Sebald <dan...@ie...>: > On 10/13/2017 07:56 AM, Eric S. Raymond wrote: > >You haven't heard from me in a couple of days because I decided to > >build a ChangeLog parser into reposurgeon and have been working on > >that. > > > >It's not an easy problem, and I have doubts that the information > >extracted that way will either cover a lot of commits or be of > >useful quality. > > I have done such a thing with the post git-converted repository already. > Code is here: > > https://sourceforge.net/p/gnuplot/patches/763/ > > I believe it is rather accurate, more accurate than the method you are > describing which attempts to pull info from the ChangeLog entries alone. > > The methodology in the link above has the added, important detail > (particular to the gnuplot project) of working from the diff hunks of the > ChangeLog file itself for every git-translated changeset. It hinges on the > fact that Ethan, et al. have consistently modified the proper location in > the ChangeLog file for every CVS modification they've made. It's that action > over the years which makes it accurate, i.e., the method traces the > incremental changes made by the maintainers. > > Parsing the ChangeLog on its own won't work very well because, as I pointed > out, the entries have been generated piecemeal. Ethan, et al. would often > make a mod that goes back several days or possibly weeks pertaining to some > checkin attributed to various contributors. Chronological order is not > maintained; the mod maintainers make can be weeks apart from the date listed > in the ChangeLog entry. > > That's why relying on the particular mod the maintainers made is better. > > The utility generates a list similar to the following: > > 87ca52fd56cf60be023642ad2664390affd5d896 2017-09-03 Daniel J Sebxxx > <xxxxxx@xxxxxx> > 778fe3d377116dc6d8c4f855372795bea5339e62 2017-09-03 Ethan A Merrxxx > <xxxxxx@xxxxxx> > c67c414e1604b20c562451ca6a2d6e9dca71561f 2017-09-02 Ethan A Merrxxx > <xxxxxx@xxxxxx> > 357f0cf68d691f3f76aade64053a6791f2e28686 2017-09-01 Bastian Maerkixxx > <xxxxxx@xxxxxx> > 984f9ad6c62f1c3ade9a41b46b5e034a2cb01c0a 2017-08-31 Martin Satuxxx > <xxxxxx@xxxxxx> > a0b903cbcae615cef9c0c0e3942c78ab150e3c82 2017-08-24 Ethan A Merrxxx > <xxxxxx@xxxxxx> > > I pointed out in a previous email that the date in the above list is of no > value for the git repository. There will not be good correlation between > the date above and the actual date of the git changeset associated with > 87ca52fd56cf60be023642ad2664390affd5d896, for example. All that needs to be > done is some automated git command that will associate the author > information in the above list with the gitID changeset in the list. I see your point. Unfortunately, I can't integrate this technique into my conversion pipeline as-is, because (a) it relies, as you note, on a regularity in metadata specific to your project, and (b) My tools know nothing about SHA-1 IDs! Those are specific to an instantiated repository's hash chains. However, there may be a way around this. Can you generate your report with the hashes replaced by action stamps? An action-stamp is a kind of commit ID my tools understand. It has this form: <committer-id!yyyy-mm-ddThh:mm:ssZ> That is, it consists of a committer ID email address, followed by an RFC3339-format representation of the commit date. These are not necessarily unique per repository, but they usually are. This command will generate an action stamp from a SHA-1 ID: git log --format='<%cE!%cIZ>' -1 $1 -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Daniel J S. <dan...@ie...> - 2017-10-13 17:54:23
|
On 10/13/2017 11:40 AM, Eric S. Raymond wrote: > Alas, automated ChangeLog mining to fill in author slots doesn't work > well enough to be usable. > > The algorithm I ended up implementing digests the ChangeLogs into a > set of tuples each consiring of a date stamp (with no time part), a > contributor ID, and a set of paths. It then walks through the entry > list, looking for commits that match the path set. Then it filters > for commits close in time to the entry timestamp. > > "Close" is by default 12 hours to either side of the commit stamp. The > window can be extended on the future side to allow for delay in > merging patches. > > $ reposurgeon "read ." "changelogs" > reposurgeon: Fills 123 of 49627 authorship slots from 4915 ChangeLog entries > $ reposurgeon "read ." "changelogs 36" > reposurgeon: Fills 468 of 49627 authorship slots from 4915 ChangeLog entries > $ reposurgeon "read ." "changelogs 72" > reposurgeon: Fills 498 of 49627 authorship slots from 4915 ChangeLog entries > $ reposurgeon "read ." "changelogs 120" > reposurgeon: Fills 515 of 49627 authorship slots from 4915 ChangeLog entries > > The argument of changelogs is a count of hours to extend the > closeness window by. As expected, increasing the window yields > more matches. > > However, even with a very long window the match rate never goes above > 1% of commits. That is noise level. > > I think the reason is hinted at by the large disparity (about 10:1) > between commit cliques and ChangeLog entries. What this tells us is > that a typical ChangeLog entry corresponds not to one commit clique > but to several. There's no algorithmic way to know what the boundaries > are. The necessary information is in the ChangeLog diff hunks of the repository itself. > If such annotations are going to be made they will need a human eye > and hand comparing at each of 4916 ChangeLog entries against the > commit history. No, see the automated utility I mentioned, which I believe is accurate and reasonable. The list ends up being 6132 entries. That suggests that on average 20% of the time the maintainers went back to a previous ChangeLog entry making changes. Dan |
|
From: Daniel J S. <dan...@ie...> - 2017-10-13 17:08:45
|
On 10/13/2017 07:56 AM, Eric S. Raymond wrote: > You haven't heard from me in a couple of days because I decided to > build a ChangeLog parser into reposurgeon and have been working on > that. > > It's not an easy problem, and I have doubts that the information > extracted that way will either cover a lot of commits or be of > useful quality. I have done such a thing with the post git-converted repository already. Code is here: https://sourceforge.net/p/gnuplot/patches/763/ I believe it is rather accurate, more accurate than the method you are describing which attempts to pull info from the ChangeLog entries alone. The methodology in the link above has the added, important detail (particular to the gnuplot project) of working from the diff hunks of the ChangeLog file itself for every git-translated changeset. It hinges on the fact that Ethan, et al. have consistently modified the proper location in the ChangeLog file for every CVS modification they've made. It's that action over the years which makes it accurate, i.e., the method traces the incremental changes made by the maintainers. Parsing the ChangeLog on its own won't work very well because, as I pointed out, the entries have been generated piecemeal. Ethan, et al. would often make a mod that goes back several days or possibly weeks pertaining to some checkin attributed to various contributors. Chronological order is not maintained; the mod maintainers make can be weeks apart from the date listed in the ChangeLog entry. That's why relying on the particular mod the maintainers made is better. The utility generates a list similar to the following: 87ca52fd56cf60be023642ad2664390affd5d896 2017-09-03 Daniel J Sebxxx <xxxxxx@xxxxxx> 778fe3d377116dc6d8c4f855372795bea5339e62 2017-09-03 Ethan A Merrxxx <xxxxxx@xxxxxx> c67c414e1604b20c562451ca6a2d6e9dca71561f 2017-09-02 Ethan A Merrxxx <xxxxxx@xxxxxx> 357f0cf68d691f3f76aade64053a6791f2e28686 2017-09-01 Bastian Maerkixxx <xxxxxx@xxxxxx> 984f9ad6c62f1c3ade9a41b46b5e034a2cb01c0a 2017-08-31 Martin Satuxxx <xxxxxx@xxxxxx> a0b903cbcae615cef9c0c0e3942c78ab150e3c82 2017-08-24 Ethan A Merrxxx <xxxxxx@xxxxxx> I pointed out in a previous email that the date in the above list is of no value for the git repository. There will not be good correlation between the date above and the actual date of the git changeset associated with 87ca52fd56cf60be023642ad2664390affd5d896, for example. All that needs to be done is some automated git command that will associate the author information in the above list with the gitID changeset in the list. > By the way, while working on this I found a bad date 2206-07-27 > in term/lua/ChangeLog. That probably wants to be 2006; somebody > should fix it. Ethan will have to fix that, thanks. Dan |
|
From: <es...@th...> - 2017-10-13 16:40:54
|
Alas, automated ChangeLog mining to fill in author slots doesn't work well enough to be usable. The algorithm I ended up implementing digests the ChangeLogs into a set of tuples each consiring of a date stamp (with no time part), a contributor ID, and a set of paths. It then walks through the entry list, looking for commits that match the path set. Then it filters for commits close in time to the entry timestamp. "Close" is by default 12 hours to either side of the commit stamp. The window can be extended on the future side to allow for delay in merging patches. $ reposurgeon "read ." "changelogs" reposurgeon: Fills 123 of 49627 authorship slots from 4915 ChangeLog entries $ reposurgeon "read ." "changelogs 36" reposurgeon: Fills 468 of 49627 authorship slots from 4915 ChangeLog entries $ reposurgeon "read ." "changelogs 72" reposurgeon: Fills 498 of 49627 authorship slots from 4915 ChangeLog entries $ reposurgeon "read ." "changelogs 120" reposurgeon: Fills 515 of 49627 authorship slots from 4915 ChangeLog entries The argument of changelogs is a count of hours to extend the closeness window by. As expected, increasing the window yields more matches. However, even with a very long window the match rate never goes above 1% of commits. That is noise level. I think the reason is hinted at by the large disparity (about 10:1) between commit cliques and ChangeLog entries. What this tells us is that a typical ChangeLog entry corresponds not to one commit clique but to several. There's no algorithmic way to know what the boundaries are. If such annotations are going to be made they will need a human eye and hand comparing at each of 4916 ChangeLog entries against the commit history. That could be done with reposurgeon; I ran some numbers and it's probably about 80 hours of hand-work. No thanks... -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> A true libertarian supports free enterprise, opposes big business; supports local self-government, opposes the nation-state; supports the National Rifle Association, opposes the Pentagon. -- Edward Abbey |
|
From: <es...@th...> - 2017-10-13 12:56:43
|
You haven't heard from me in a couple of days because I decided to
build a ChangeLog parser into reposurgeon and have been working on
that.
It's not an easy problem, and I have doubts that the information
extracted that way will either cover a lot of commits or be of
useful quality.
Here's what my documentation now says about the feature:
Mine the latest version of every ChangeLog file for authorship data.
Assume such files have basenames beginning with 'ChangeLog', and that
they are in the format used by FSF projects: entry header lines begin
with YYYY-MM-DD and are followed by a fullname/address, entry file
path lists in the text are delimited by '*' and ':' and each list
always begins first on a line after whitespace.
The logic attempts to match commits by path manifest and date; the
ChangeLog from which an entry derives is added to its manifest.
Because of time zone issues, each entry date is allowed to match
commit dates 12 hours earlier or later in UTC.
To allow for lag in applying inbound patches, the forward window can be
extended by a length of time given as a command-line argument in hours;
the default is 0. Note that setting this high is likely to cause false
matches, especially on single-file commits.
When an commit matches an entry, the attribution and the date from the
entry are copied to the author list of the commit. This can happen
more than once per entry. Author attributions generated this way will
always have a timestamp of 00:00:00 with an offset of +0000.
These attributions are marginally better than nothing but should not
be greatly trusted.
Alas, the problems here are fundamental, stemming from inadequate data
representations, and cannot be fixed with clever programming.
I have written the ChangeLog parser stage and am working on the
logic to match ChangeLog entries to commits.
By the way, while working on this I found a bad date 2206-07-27
in term/lua/ChangeLog. That probably wants to be 2006; somebody
should fix it.
--
<a href="http://www.catb.org/~esr/">Eric S. Raymond</a>
A true libertarian supports free enterprise, opposes big business;
supports local self-government, opposes the nation-state; supports the
National Rifle Association, opposes the Pentagon. -- Edward Abbey
|
|
From: Achim G. <Str...@ne...> - 2017-10-12 18:04:48
|
sfeam via gnuplot-beta writes: > I will make that change. No promises that another wrinkle will not > appear. Well, once I see it appear in the repo I'll do a new build and run our plotting scripts @work against it. I had to rollback to 5.0.7 pretty quickly so I can't really say if there were no other problems, but nothing else that was glaringly obvious. > Some of the very oldest entries in the project bug tracker > were complaints about the handling of log-scale tic placement. I can imagine. > The revision of log-scale handling in 5.0 allowed us to close them > a mere 25 years later. But even so it's evidently not yet perfect. I just needs to be good enough. :-) Regards, Achim. -- +<[Q+ Matrix-12 WAVE#46+305 Neuron microQkb Andromeda XTk Blofeld]>+ SD adaptation for Waldorf microQ V2.22R2: http://Synth.Stromeko.net/Downloads.html#WaldorfSDada |
|
From: Eric S. R. <es...@th...> - 2017-10-12 10:57:28
|
Hans-Bernhard Bröker <HBB...@t-...>: > Am 11.10.2017 um 19:01 schrieb Ethan A Merritt: > > >I do not know which email address Hans-Bernhard Broeker would prefer > > > >broeker = Hans-Bernhard Broeker <br...@ph... > ><mailto:br...@ph...>> Europe/Berlin > > My physik.rwth-aachen.de address is gone, and has been for a while. I've > kept using it as the check-in address out of tradition, mainly. > > Please use broeker .at. users.sourceforge.net, instead. Done. > > > 3. I'm not clear on what ought to be excised from the repo. HBBroeker > > > mentions dropping the faq module. Someone else (I think Mojca) had > > > a pre-conversion sript that stripped out some stuff. A set of excisions > > > you guys agree on - ideally expressed as a pre-conversion script > > >modifying the repo - is one thing we need. > > Using reposurgeon, the 'faq' module is already excluded because its > generated Makefile rule for SourceForge inputs only pulls one module at a > time, so there's nothing to excise. If anything, this module might best be > converted independently, and the resulting git repo plunked into the new > gnuplot git repo, e.g. into docs/old. OK, this can be done. I have successfully fetched the faq module > > > there were some files on ther gitspace side that weren't in CVS (probably > > > due to a botched CVS delete). This needs to be further investigated. > > I didn't mention those before because I'm convinced the reconstruction is > actually better than the original in this case. That is a happy accident. >cvsconvert gives me a state of the converted repository like this: > > > $ git status > > HEAD detached at pre-pm3d-11 > > nothing to commit, working tree clean > >Back when that branch was current, there were indeed no log-rotated >ChangeLogs yet. > > git checkout master > >makes ChangeLog.[0-6] appear. Yes, it does. It sounds, then, as though there are zero problems with the cvs-fast-export conversion itself as it is. The contributor map is complete; the author map still needs to be merged. I still need the ancient-release tarballs. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: sfeam <sf...@us...> - 2017-10-12 03:36:09
|
On Wednesday, 11 October 2017 19:11:08 Achim Gratz wrote: > > I've been building gnuplot for Cygwin from the stable branch over the > weekend in order to get the fix for the log axis regression. I have the > grid lines and tics back in my plot, but there is yet another > regression: > The 5.0 branch will put the major tics on integer exponents > of the base. > The 5.2 branch will only do that if the first defined tic > is located inside the axis range and otherwise uses the lower end of the > range to start from. Yup. Good diagnosis. There is a line of code that does exactly that. I agree this is probably not what people expect, so that line can go. > I hope this can still be fixed for 5.2.1 (fingers crossed). I will make that change. No promises that another wrinkle will not appear. Some of the very oldest entries in the project bug tracker were complaints about the handling of log-scale tic placement. The revision of log-scale handling in 5.0 allowed us to close them a mere 25 years later. But even so it's evidently not yet perfect. Ethan > Regards, > Achim. > |
|
From: Hans-Bernhard B. <HBB...@t-...> - 2017-10-11 20:10:29
|
Am 11.10.2017 um 19:01 schrieb Ethan A Merritt: > I do not know which email address Hans-Bernhard Broeker would prefer > > broeker = Hans-Bernhard Broeker <br...@ph... > <mailto:br...@ph...>> Europe/Berlin My physik.rwth-aachen.de address is gone, and has been for a while. I've kept using it as the check-in address out of tradition, mainly. Please use broeker .at. users.sourceforge.net, instead. > > 3. I'm not clear on what ought to be excised from the repo. HBBroeker > > mentions dropping the faq module. Someone else (I think Mojca) had > > a pre-conversion sript that stripped out some stuff. A set of excisions > > you guys agree on - ideally expressed as a pre-conversion script > > modifying the repo - is one thing we need. Using reposurgeon, the 'faq' module is already excluded because its generated Makefile rule for SourceForge inputs only pulls one module at a time, so there's nothing to excise. If anything, this module might best be converted independently, and the resulting git repo plunked into the new gnuplot git repo, e.g. into docs/old. > > there were some files on ther gitspace side that weren't in CVS (probably > > due to a botched CVS delete). This needs to be further investigated. I didn't mention those before because I'm convinced the reconstruction is actually better than the original in this case. There are a date range of old tags/branches (BETA_347_980818 ... BETA_349_990114, GNUPLOT_RELEASE_3_7_0, GNUPLOT_3_7_0_2 to GNUPLOT_3_7_0_3, and GNUPLOT_990126 ... GNUPLOT_990317) for which the reconstructed git repo has a subdirectory docs/ps But the comparison didn't find it in its CVS checkouts. That folder did exist between 1998-07-21 and 1999-03-18, then it was copied to docs/psdoc, and the original dropped. So the archives for these files are now in gnuplot/docs/Attic/ps/Attic/*,v Looks like CVS can't really cope with subdirectories inside an Attic. FWIW, cvs2git also rejects this: Directory cvs/gnuplot/docs/Attic/ps found within Attic; ignoring cvsconvert has effectively the same result: git tags/branches from that date range report five gitspace-only files in that folder because CVS fails to express files that should have been there. |
|
From: Eric S. R. <es...@th...> - 2017-10-11 18:24:34
|
Ethan A Merritt <sf...@us...>: > Bastian Maerkisch sent an updated list, which I copy below Saved. > I do not know which email address Hans-Bernhard Broeker would prefer I guess he'll tell us. > vanzandt = James R. Van Zandt <jr...@va...> > vanzandt = James R. Van Zandt <jr...@de...> Can someone find out which he prefers? > > Ideally, we'd enhance this in two ways. First, those of you with preferred > > email addresses should make sure this points at the one you want. Second, > > we'd add timezone offsets. > > I don't think a timezone correction is appropriate. > My experience in committing to SourceForge is that the > timestamp is recorded in UMT rather than as my local time. > So at least for my commits it would be incorrect to apply a shift > based on the original location of the author. You are right. In the conversion process he timezone is *not* used to apply a time shift. Rather, it is used to set the time zone offset that git will use for *display* of the time. This is really just a cosmetic feature; the times are stored internally as UTC and without a per-committer timezone they're displayed as UTC too. > > 2. If you want the authorship data to be right, somebody's going to have to > > write code to grovel through the ChangeLog(s) and generate > > date/author/filelist triples. (Then I'll have to write some custom Python > > to use those.) Though maybe this isn't worth it - the only > > ChangeLog I can see only seems to cover 1998-2000. > > ChangeLog 2014-08-21 - 2017-10-09 > ChangeLog.5 2014-03-15 - 2014-08-21 (yes this one is redundant) > ChangeLog.4 2011-11-22 - 2014-08-20 > ChangeLog.3 2009-10-18 - 2011-11-22 > ChangeLog.2 2006-10-01 - 2009-10-17 > ChangeLog.1 2004-04-17 - 2006-10-01 > ChangeLog.0 1998-04-09 - 2004-04-15 That's odd. The ChangeLog.[0-5] files don't show in my conversion. I will investigate. > > 3. There was some talk of gluing to the history old releases that only > > exist as tarballs. That can be done, but I need those tarballs. > > I will Email a copy privately. OK. -- <a href="http://www.catb.org/~esr/">Eric S. Raymond</a> My work is funded by the Internet Civil Engineering Institute: https://icei.org Please visit their site and donate: the civilization you save might be your own. |
|
From: Achim G. <Str...@ne...> - 2017-10-11 17:11:36
|
I've been building gnuplot for Cygwin from the stable branch over the weekend in order to get the fix for the log axis regression. I have the grid lines and tics back in my plot, but there is yet another regression: The 5.0 branch will put the major tics on integer exponents of the base. The 5.2 branch will only do that if the first defined tic is located inside the axis range and otherwise uses the lower end of the range to start from. This is interesting and maybe even useful sometimes, but it breaks a lot of plotting scripts, so it shouldn't do this by default I think. To reproduce: --8<---------------cut here---------------start------------->8--- set log x 10 set mxtics 10 set xtics 10 set xrange [pi:1e4*pi] plot sin(x) --8<---------------cut here---------------end--------------->8--- --8<---------------cut here---------------start------------->8--- set log x 10 set mxtics 10 set xtics 3,10 set xrange [pi:1e4*pi] plot sin(x) --8<---------------cut here---------------end--------------->8--- --8<---------------cut here---------------start------------->8--- set log x 10 set mxtics 10 set xtics 4,10 set xrange [pi:1e4*pi] plot sin(x) --8<---------------cut here---------------end--------------->8--- I hope this can still be fixed for 5.2.1 (fingers crossed). Regards, Achim. -- +<[Q+ Matrix-12 WAVE#46+305 Neuron microQkb Andromeda XTk Blofeld]>+ Waldorf MIDI Implementation & additional documentation: http://Synth.Stromeko.net/Downloads.html#WaldorfDocs |
|
From: Ethan A M. <sf...@us...> - 2017-10-11 17:04:12
|
On Wednesday, 11 October, 2017 09:16:39 Eric S. Raymond wrote: > sfeam <sf...@us...>: > > It looks to me that there is general consensus that > > > > - we should move to git, > > > > - that Eric Raymond's toolset is the best way to do that, and that > > > > - if Eric himself is willing to guide the conversion it has the greatest chance of success. > > > > So let's do that. > > This is a good time for me to do it. NTPsec 1.0 just shipped and our PM said > he didn't want to see any commits for a week. :-) > > > Eric - several people responded with bits and pieces of meta-info that > > you asked for. Do you have everything you need? > > Not yet. > > 1. Here's the committer map: > > janert = Philipp K. Janert <ja...@us...> > uid225733 = Ethan Merritt <merritt@u.washington.edu> > uid93776 = Ethan Merritt <merritt@u.washington.edu> > mikulik = Petr Mikulik <mi...@ph...> > markisch = Bastian Maerkisch <bma...@we...> > tlecomte = Timothee Lecomte <tim...@en...> > persquare = Per Persson <per...@us...> > amai = Alexander Mai <am...@us...> > lhecking = Lars Hecking <lhe...@us...> > cgaylord = Clark Gaylord <cga...@vt...> > uid26705 = Petr Mikulik <mi...@ph...> > juhaszp = Peter Juhasz <ju...@us...> > vanzandt = James R. Van Zandt <van...@us...> > sfeam = Ethan Merritt <merritt@u.washington.edu> > lodewyck = J´erˆome Lodewyck <lod...@us...> > joze = Joze Duhovnik <jo...@us...> ^^^^ No. this one is not correct. "joze" refers to Johannes Zellner Bastian Maerkisch sent an updated list, which I copy below I do not know which email address Hans-Bernhard Broeker would prefer amai = Alexander Mai <st0...@hr...> Europe/Berlin broeker = Hans-Bernhard Broeker <br...@ph...> Europe/Berlin cgaylord = Clark Gaylord <cga...@vt...> US/Eastern janert = Philipp K. Janert <ja...@ie...> US/Pacific joze = Johannes Zellner <joh...@ze...> juhaszp = Peter Juhasz <ju...@us...> lhecking = Lars Hecking <lhe...@us...> lhecking = Lars Hecking <lhe...@nm...> Europe/Dublin lodewyck = Jérôme Lodewyck <lod...@us...> markisch = Bastian Maerkisch <bma...@we...> Europe/Berlin mikulik = Petr Mikulik <mi...@ph...> Europe/Prague persquare = Per Persson <per...@ma...> sfeam = Ethan A Merritt <merritt@u.washington.edu> US/Pacific tlecomte = Timothee Lecomte <tim...@en...> Europe/Paris vanzandt = James R. Van Zandt <jr...@va...> vanzandt = James R. Van Zandt <jr...@de...> uid26705 = Petr Mikulik <mi...@ph...> uid93776 = Ethan A Merritt <merritt@u.washington.edu> > Ideally, we'd enhance this in two ways. First, those of you with preferred > email addresses should make sure this points at the one you want. Second, > we'd add timezone offsets. I don't think a timezone correction is appropriate. My experience in committing to SourceForge is that the timestamp is recorded in UMT rather than as my local time. So at least for my commits it would be incorrect to apply a shift based on the original location of the author. > 2. If you want the authorship data to be right, somebody's going to have to > write code to grovel through the ChangeLog(s) and generate > date/author/filelist triples. (Then I'll have to write some custom Python > to use those.) Though maybe this isn't worth it - the only > ChangeLog I can see only seems to cover 1998-2000. ChangeLog 2014-08-21 - 2017-10-09 ChangeLog.5 2014-03-15 - 2014-08-21 (yes this one is redundant) ChangeLog.4 2011-11-22 - 2014-08-20 ChangeLog.3 2009-10-18 - 2011-11-22 ChangeLog.2 2006-10-01 - 2009-10-17 ChangeLog.1 2004-04-17 - 2006-10-01 ChangeLog.0 1998-04-09 - 2004-04-15 > 3. There was some talk of gluing to the history old releases that only > exist as tarballs. That can be done, but I need those tarballs. I will Email a copy privately. I leave it to others to respond to points 3 + 4 Ethan > 3. I'm not clear on what ought to be excised from the repo. HBBroeker > mentions dropping the faq module. Someone else (I think Mojca) had > a pre-conversion sript that stripped out some stuff. A set of excisions > you guys agree on - ideally expressed as a pre-conversion script modifying > the repo - is one thing we need. > > 4. The following paragraph from HBBroeker worries me a little: > > > I've tried out reposurgen on our repository (with the "faq" module > > removed), and it seems to work almost perfectly. "make allcompare" > > showed nothing missing from the git version. The original import > > pseudo-branch GNUPLOT_BETA is, correctly, dropped, all other tags and > > branch tips compare equal, and .cvsignore files are also taken care of. > > I didn't use reposurgeon, I used cvsconvert. This is a wrapper script > in the reposurgeon distribution that also uses cvs-fast-export as an > engine, but is specialized for CVS and does more detailed correctness > checking than allcompare. What worries me is that *my* conversion didn't > look quite so smooth - it was not clear that GNUPLOT_BETA was dropped, and > there were some files on ther gitspace side that weren't in CVS (probably > due to a botched CVS delete). This needs to be further investigated. > > > Do I understand correctly that the git repository can be created > > initially anywhere that you find convenient and then replicated to > > SourceForge afterwards? > > That understanding is correct. > > > Should we set a freeze date for the existing CVS source [*]? > > That's not important yet. Once I have the conversion process > scripted, you can basically choose any time to cut over and it > will all get done in a time on the close order of two hours. > > > Anything else that needs to be done in preparation? > > For the conversion itself, no. You guys need to make a hosting site > decision, then I need to have push and force-push privileges so I can drop > the git repo in place. > |
|
From: Daniel J S. <dan...@ie...> - 2017-10-11 15:50:00
|
On 10/11/2017 08:16 AM, Eric S. Raymond wrote: > sfeam <sf...@us...>: >> It looks to me that there is general consensus that >> >> - we should move to git, >> >> - that Eric Raymond's toolset is the best way to do that, and that >> >> - if Eric himself is willing to guide the conversion it has the greatest chance of success. >> >> So let's do that. > > This is a good time for me to do it. NTPsec 1.0 just shipped and our PM said > he didn't want to see any commits for a week. :-) > >> Eric - several people responded with bits and pieces of meta-info that >> you asked for. Do you have everything you need? > > Not yet. > > 1. Here's the committer map: > [snip] > > Ideally, we'd enhance this in two ways. First, those of you with preferred > email addresses should make sure this points at the one you want. Second, > we'd add timezone offsets. I didn't know timezone offsets for individuals is possible. How is that advantageous in terms of git usage? Is it that when git displays stuff to the user it adjusts date/time? Does that happen in viewers like gitg, "git gui", etc.? > 2. If you want the authorship data to be right, somebody's going to have to > write code to grovel through the ChangeLog(s) and generate > date/author/filelist triples. (Then I'll have to write some custom Python > to use those.) Though maybe this isn't worth it - the only > ChangeLog I can see only seems to cover 1998-2000. I've written a utility that constructs the list you are looking for by examining all the ChangeLog diff hunks after the repository is translated. The code is here https://sourceforge.net/p/gnuplot/patches/763/ and the algorithm is based on searching the diff hunks for the first changed line (i.e., "+" in the first character) after the "+++" of the headers. After finding that first change, it searches backward for the author info of the ChangeLog on that particular version. Providing enough context lines in the diff hunks should allow capturing that authorship line. Example usage is: git log -p --unified=50 ChangeLog > ChangeLog.diff git_changelog_author ChangeLog.diff > author.txt (I suppose I could have written the utility to allow pipe redirection to skip the intermediate file.) Examining the ChangeLog.diff file in an editor should go a long way to understanding how the algorithm works. The output is gitID/Date/Author/Address separated by space character (allowing the Author to have spaces as well). The date information is of no use to reconstruction in the git repository--as I explained in a previous email, any particular ChangeLog entry can be constructed piecemeal over the course of a day. (But the utility should properly assign the same author info to all the incremental changes.) I would think a python script could step through the list and retroactively call some git command to correct the authorship. That would be nice. I hope it doesn't make an entry in the git repository for every authorship change somewhere, though, as there are over 6000 authorship lines. (BTW, those changesets that had no modification to ChangeLog go unrecognized by the utility, but that's fine.) I notice in the list of authorship there are some variations on names/address. For example, Ethan is most often E. M., but often it appears E. A. M. Do we want to make names consistent? Those sort of variations are probably going to happen going forward in git as well, seeing as any particular person might do mods from different computers with slightly different user info configuration. The list of files would come from the git repository changeset itself, but honestly I don't think that's of any benefit without any comments to go along with it. Two reasons: 1) If it is just a list of files, that's already gotten from something like gitg viewer which lists all the files modified in a particular changeset, and 2) conveniently the name "ChangeLog" often being the first in the list puts the diff hunk for the ChangeLog right after the description so we sort of have the info format we want already, e.g., from gitg: broeker <broeker> 10/06/2017 06:35:09 PM +0000 Use if available. Provide centralized fall-back of WEXITSTATUS if does not supply it. Expand all 0a3035e39a1f9402a474d0ea9300deac4cd9ef33 .... .... ▼15 ChangeLog .... .... @@ -1,3 +1,14 @@ 1 + 2017-10-06 Xxxxxx Xxxxxx <xxxxxx@xxxxxx> 2 + 3 + * src/command.c: Move WEXITSTATUS fall-back definition away from here. 4 + 5 + * src/syscfg.h: Include <sys/wait.h>, if it exists. 6 + (WEXITSTATUS): Provide fall-back definition, if none in 7 + <sys/wait.h>. Move MS Windows specific replacement from command.c 8 + to here. 9 + 10 + * configure.ac: Add call to AC_HEADER_SYS_WAIT 11 + 1 12 2017-10-06 Xxxxxx Xxxxxx <xxxxxx@xxxxxx> 2 13 3 14 * config/mingw/Makefile: Add helpfiles to "all" target, including The above works for me, as far as retrieving a more detailed description of the changeset, i.e., no need to consult historical ChangeLog on this one. A discussion worth having (gnuplot group) is what to do with the ChangeLog going forward. Rather than ChangeLog ChangeLog.0 ChangeLog.1 ChangeLog.2 ChangeLog.3 ChangeLog.4 ChangeLog.5 it might be nice to put all those in some file ChangeLog_CVS or ChangeLog_historical--something that indicates this is really no longer an active file and all such info (similarly formatted descriptions) will be in the git changeset comment going forward. Also, leaving said file in the root directory is best because I often use "grep */*" to quickly search for information in the source tree and prefer not all sorts of entries from the ChangeLog appearing in the search list. > 3. There was some talk of gluing to the history old releases that only > exist as tarballs. That can be done, but I need those tarballs. If you were to first create a draft posted somewhere, we could clone the repository and test whether the tags for a particular version will retrieve code that matches the tarball. If so, then we don't need tar balls within the repository, do we (group)? Dan |