Wednesday, July 16, 2014

Monday, July 14, 2014

Vacation thoughts - old bug dead in new design

Too busy doing vacation stuff to think much about PPQT but one idea did occur in a quiet moment: that under the new design I sketched in the prior post, an old and difficult bug will no longer matter.

PPQT uses a list of QTextCursor objects to track the start of each logical page of scanned text. This is an important function. It enables me to display the correct scan image as the user moves the cursor. It enables auto-generation of page-marker text, as for HTML conversion.

When a book is first opened, the list is set up by recognizing the PGDP page separator lines. Later, the user deletes those lines, but a QTextCursor doesn't care, it just keeps pointing between character A, at the end of page A, and character B, the first letter of page B, and is automatically updated by QTextEdit as characters are added and deleted (and moved, which is a sequence of delete and add).

It all works perfectly except for a major bug. If a text insertion or deletion spans the position of a QTextCursor, the position is changed to the end of the inserted or deleted text. This is arguable correct. But if the user then Undoes that change, the cursor's position is not restored. QTextCursor values are not saved on the undo/redo stack. (Here's the bug report I filed.)

Changing text that spans a page boundary is unusual in normal editing but common in one case: paragraph reflow for the ASCII conversion. A paragraph may well span a page boundary. In V.1, the reflow logic has a lot of code to notice this case and preserve the page boundary cursor, so that after all the lines of the old paragraph are replaced in one operation by the lines of the reformatted paragraph, the boundary cursor is fixed up to point between the same characters A and B as before.

But if the user then undoes the reflow, the boundary cursor jumps to the end of the restored, original paragraph and points to the wrong place. The page image no longer flips when it should. Users are encouraged to test-reflow paragraphs, tables, poems, etc., and to ctl-Z the reflow if the result is not good. Which pretty much guarantees that many page boundary cursors will not survive the ASCII conversion step (or the HTML conversion, which is similar).

I've put some thought into this and never saw a way to fix it. I've thought about, and meant to experiment with, customizing the Undo/Redo classes, possibly storing an affected cursor's position in the stacked element (sorry, can't be bothered to look up its class name just now) and fixing it during a redo. But there are a lot of questions, such as, when does QTextEdit update the affected cursors' positions, before or after stacking an undo element?

Then I realized, under the translator scheme I outlined previously, the whole issue goes away. ASCII reflow is part of "translating" from DP markup to a different markup, the PG etext format. Under the new design, this will never be done ad-hoc or section by section. It will always be done as a complete translation of the DP text producing a new file in a new edit tab. The DP text is not affected and its page cursors are still good. The only kind of "undo" the user can have, is to just close the new book without saving it, and translate again.

So I can stop worrying about maintaining page cursors over reflow. I do have to worry about passing page boundary information to a "translator" function, so the new, reformatted book can have its own unique table of page boundary cursors, but I think I know how to do that.

Not sorry to see an old bug die.

Sunday, June 29, 2014

Design Breakthrough, and July Off

A couple of posts back I agonized about how PPQT should relate to the new, competing markup styles. Today I figured out how to handle this.

The recommended and expected workflow will be that the PPer first brings the text to a state of completion with respect to the DP formatting guidelines. Heal all page breaks, fix all typos, deal with all proofer notes, move and renumber footnotes, process all gutcheck/bookloupe diagnostics. Delete all page separator lines. Also, use a new dialog to set metadata for the book: title string, author string at least.

At this point the PPer will choose File>Generate... and get a dialog with a menu of possible "translators", each able to translate a DP text into some other markup. From the user's point of view, now something magical happens: after a moment of high CPU usage, a new book appears. Its filename is "Untitled-n" (same File>New makes) and its contents are the translated text of the starting book. The user then uses File > Save As to save it under some appropriate name and suffix.

The magic that happens under the covers is this. Code that I will write parses the input text, extracting all possible information out of the DP markup conventions. Using a simple and rather elegant API that I have in mind, it will pass these data to the chosen translator. This API will make it so stupidly simple to generate a translated document that anybody could write one.

I will write a plain-text translator and an HTML translator. These can be used as models. Anybody else who wants to write a translator is welcome to do so and to issue a pull request. A translator for fp; a translator for fpgen; a translator for XML; a translator for Markdown or reStructured text or whatever, I don't care, let a thousand goddam flowers bloom.

Holidays

In a couple of days we are off for a month in Scandinavia: Copenhagen, Bergen, Trondheim, Stockholm. In the unlikely event that any reader of this blog wants to follow our very mild adventures, you can do so in our Scandinavian travel blog.

Posting in this blog will resume in August. See you then.

Thursday, June 26, 2014

Doing It "Pythonically" (with added thought)

I'm starting work on the Find module. One of the features of that is that each of the input fields—the Find text and each of three independent Replace texts—has a memory pop-up: a button that pops up a menu of the ten most recent strings used in that field. It's a very nice help when you are alternating between two or three complex regex searches. (And stolen from Guiguts and BBEdit.)

The details of this widget were of course encapsulated in a class; for V.2 it is class RecallMenuButton. In V.1 this was a QComboBox because that is easy to present. I maintained the recent strings in a QStringList, and any time the list was updated the widget could reload itself in one call, self.insertItems(list). However, I wanted it to look like a square button, not a list, so I had it set its own max width to 29px. Under Windows that didn't work with the default style, so I had to set it to a non-native style "CleanLooks", and that no longer exists.

Anyway for V.2 I am using a "command button" which is a button that has an associated menu. I maintain a python list of QActions, one for each string. When the menu for the button emits the aboutToShow signal, I clear the menu and add the list of actions to it.

So I was starting to code the remember() method of this class, which takes a string and, if it is in the list now, deletes it; then adds it to the front of the list. But some considerations:

  • The string might very likely be in the list already as the first item, because it will be frequent to Find or Replace the same string over and over.
  • The list might be empty, at least the first time.
  • The string might not be in the list at all.
  • The list should never exceed MAX_STRINGS in length.

So the V.1 code is about a dozen lines, with a Fortran-like loop over the list to find and remove the string if it exists. But I'd like to do it more "pythonically" this time. I came up with this:

    def remember(self, string):
        new_stack = [act for act in self.string_stack if string != act.text()]
        new_stack[0:0] = QAction(string)
        self.string_stack = new_stack[0:self.MAX_STRINGS]

After composing which I said,

Now, this does not try to short-cut the presumptively common case of the string already being on top of the stack. If the added string is at the front of the stack it will be removed and added back in the same position. Is this a waste of time? Yes, but the test for that special case would look like:

    def remember(self, string):
        if len(self.string_stack):
            if string == self.string_stack[0].text() :
                return

...and that much extra code would surely waste as much time as it saved, or nearly. Or would it? Hmmm.

Edit: no, it would not be worth it and here's why. The Find UI is only going to "remember" a find or replace string if that string is manually edited by the user. When the Find or Replace line-edit field is filled by the program (from the remembered-strings popup menu or from a user-loadable macro button) a flag is cleared. Said flag is only set when the user edits in the field (the textChanged signal). When the field is actually used (Find or Replace action done), the flag is tested and remember() is called only if the flag is set. Thus remember() is only ever called for strings that the user has entered or altered. This greatly reduces both the number of calls to remember() and the chances that a remembered string already exists at any position in the stack. So there is no point in guarding against the input string being at stack[0].

Tuesday, June 24, 2014

Notes functional but ugly

I finished and pushed the notesview module and a cursory unit test for it. I suppose I should test it more thoroughly but I exercised all its features including errors, so pffft.

Part of this job was a bit of refactoring back in the editview module. It has a line-number widget that the user can type into. On hitting Enter, the editor jumps to that line number. Similarly in the image name widget: type '125' and hit enter to jump to the page for image 125.png. These were coded as slots to receive the returnPressed signals from those widgets.

Well, about the only feature the Notes panel has, other than being a plain text editor, is that you can hit shift-ctl-M to record the current edit line number in your notes as {nnn}. And then you can put the cursor in or by a line number and hit ctl-M to jump the editor to that line.

n.b. it started out as [shift-]ctl-L, mnemonic for "line", but that turned out to have a dedicated use in Ubuntu. Locking the screen, I seem to recall; embarrassing thing to have happen when you only meant to note a line number.

Also, shift-ctl-P records the current page's image name as [xxx] in your notes, and you can put the cursor in or near a page name and it ctl-P to jump to that page.

n.b. yes, this is the accelerator for Print and Print Setup in many apps. PPQT doesn't support printing, so too bad.

Jumping to lines or pages of course necessitates that the Notes panel code interact with the Edit panel code. And only belatedly did it occur to me, that the ctl-M action was identical to hitting Enter in the line number widget box, and ctl-P was the same as hitting Enter in the image name widget box. So I had to go into the editor and factor out the guts of each action from the signal-slot it was buried in, make them separate and public methods, so the Notes panel could call them.

UI issues

Both the edit widget and the notes are derivatives of QPlainTextEdit. However they behave differently, and not for any reason I can find.

The editview editor displays selected text under the default highlight color, a nice lemon yellow. Then when you focus out of it (click in some other widget), the selection highlight changes to medium gray.

The noteview editor does not do this. It always displays selected text using the mid-gray "inactive" highlight color—even when it definitely has the focus and should be active. More puzzling, I put code in the focusInEvent() member to get its QPalette and print out the "current color group" and the name of the highlight color. The group was "Active" and the color was the same (#fbed73) lemon-yellow as the other editor. But the actual display is still gray. So this is baffling and I have posted a query at the Qt forums to see if I can get some help.

Meanwhile, on to experimenting with using the Qt Designer to lay out the complex Find panel. Probably won't work, but maybe.

Sunday, June 22, 2014

More Thoughts on Pandoc, etc. [Updated]

Some additional considerations that have occurred to me since writing the previous post.

Plain Text Output

One output format is absolutely required of the Post-Processor: the complete book as a plain text file. Formerly this had to be ASCII, not even Latin-1, with accented characters in an expanded format like [:u] or [c~]. Nowadays, PPers usually provide the ASCII, plus a Latin-1 or UTF-8 version with accented characters in place.

Regardless of the encoding, the plain etext has no formatting markup. It is simply the text with headings set off by newlines, paragraphs wrapped to a 72-character margin, and any other formatting, like poetry, tables, or centered text implemented with spaces and newlines.

PPQT V.1, like Guiguts before it, does quite a nice job of converting DPM to plain etext. I took considerable pride in implementing the Knuth-Pratt algorithm for optimal paragraph reflow, as a point of differentiation from Guiguts. And PPQT reflows tables (coded with the unique PPQT markup, of which more later) nicely also.

It occurred to me to wonder how well Markdown or any other markup accepted by Pandoc did at this task. And to my surprise, it appears that none of them do it at all!

The venerable Markdown for example plainly states that "Markdown is a text-to-HTML conversion tool." Not a plain-text generating tool, an HTML one, which means it hands off all responsibility for paragraph reflow to the web browser. Similarly AsciiDoc and reStructuredText mention only output to HTML, PDF, EPub and the like. (Well, rST mentions output to Python Docstrings, but doesn't say whether it reflows paragraphs for them.)

It seems quite likely—although I would love to be corrected on this!—that it is not possible to go into Pandoc with any markup and come out with a plain UTF-8 text file acceptable to Project Gutenberg!

Edit: According to someone on the Pandoc mailing list, there is indeed a plain-text output "writer", and pandoc -t plain my-input.txt should produce what I want. I haven't installed an actual pandoc so can't try this, for example I don't know what the paragraph reflow is like, or whether there is any way to control widths. So still not certain if plain text is really feasible from pandoc. To be investigated later this year.

Further edit: my innocent question on the Pandoc list has produced this interesting thread with some knowledgeable comments about the history of PG and its format (esp. its ambiguities, which make it very hard to back-convert PG to some markup, one example, use of CAPS for emphasis), and this from John MacFarlane (Pandoc author): "I think a pg writer is a nice idea. It would be fairly easy, I think, to do.... [it] would involve a new option and a few different behaviors." From this I deduce that the existing "plain" output mode is not fully PG-compliant.

And another edit: On the same message thread linked above, John MacFarlane (Pandoc author) now says, "I've started a gutenberg branch on github. It should be fairly easy to add a writer that uses PG conventions." A Fred Zimmerman, presumably a PG or DP contributor? adds "i'm very interested in the gutenberg branch -- great idea." So this may develop into something good, and soon!

This considerably reduces the value of Pandoc to DP. The plain etext is not a negotiable requirement. Not surprisingly the Python programs that process fpn and fpgen do promise plain-text outputs in addition to Epub, HTML, etc.

The question now is: should PPQT V.2 continue the ability to convert DPM to etext? There's a fair amount of code and GUI widgetry behind it. Or should I just assume everyone will be converting to fp[ge]n markup and getting their etext from the batch programs that support those markups?

Translation UI

The relationship between PPQT and the three competing markups (DPM, fpn, fpgen) is quite unclear to me, as should be apparent. I'm thinking I badly need to know what the potential user community actually needs—indeed, if a user community actually exists at all!. If I don't get some useful comments on this blog, I need to go to the forums and make a nuisance of myself to get some answers.

But supposing a community exists, here's what I think I know. First, DPM will not go away. What I'm calling DPM is the sum of the rules in the Distributed Proofreaders Formatting Guidelines. It is deeply embedded in the whole DP infrastructure. I don't think DP will ever rewrite their guidelines to make the volunteer proofers in the Formatting rounds insert fp[ge]n syntax instead.

If I'm right about that, then DPM is what the PPers will continue to receive as their input. The initial stages of PP work—fixing up things separated by page breaks, double-checking bold and italic markups, renumbering and moving footnotes, and running spellcheck—will be done in the context of a DPM text.

Then, most likely, the PPer will want to do a one-time bulk conversion to one of the fp[ge]n markups. This, PPQT could facilitate, in the following ways.

First, provide "File > Export to" command options. As I noted in the prior post, it would not be difficult to convert DPM to either fp* markup. This command would act like File > Save As, bringing up a file-save dialog to pick a name, and writing a new or replacement file consisting of the active (DPM) document translated to another markup: mybook.txt is saved as mybook.fpn.

Second, have a way the user can opt to save metadata with the translated file, mybook.fpn.meta. The meta file would include the pointers to where page boundaries were (adjusted to still be accurate in the translated source file), as well as the notes, bookmarks, vocabulary etc.

This pretty much means that you could now open the fp* markup file and still have your scan images, your notes, word and character tables... not sure about the footnotes. But you could go on editing as before.

Saturday, June 21, 2014

Looking ahead to Pandoc

This is to order my thoughts about output formats and markup systems, and to gather links to these in one convenient place.

History

PPQT is intended to support the work of volunteers finishing etexts for Distributed Proofreaders, aka PGDP. PGDP was one of the first "crowd-sourced" volunteer sites on the internet, organizing thousands of volunteers to find the typos in OCR images of public-domain texts, one page at a time. At the end of the process, a different set of volunteers, the "post-processors" or PPers, have the job of splicing together the individually-proofed pages of each book to make one smooth etext. That's the task that PPQT aimed to assist.

The original PGDP workflow ended with an ASCII etext, no more. There are hundreds (thousands?) of PGDP-proofed etexts at Project Gutenberg. By 2002 or so, most PPers also prepared HTML versions of their texts. And in recent years there's been demand for other formats such as EPUB.

Markup Systems

A text passing through PGDP gets formatted with a particular markup style documented in the Formatting Guidelines. Although PGDP did not label the guidelines as a "markup system" that is what they constitute: a set of rules for representing a book's typography and layout in a plain text document. PGDP never gave their markup system a catchy name; let's call it DPM.

dpm

DPM can be compared to other plain-text markups such as Markdown and reStructured Text. It comes off quite well in these comparisons. The other markups were devised by (mostly) programmers for use in (mostly) documenting code, they don't support typography beyond emphasis, and layout beyond code-blocks. Some of the things that DPM supports and others do not include footnotes, poetry (in the sense of being able to specify line breaks and indentation), and simple right-alignment of text, as in a citation within a block quote.

fpn

In recent years, PGDP volunteers motivated in part by the need to auto-convert etexts to new formats such as EPUB (and in part, I'm sure, by simple N.I.H. syndrome), have devised new markup styles. One is fpn devised by Robert Frank (rfrank at PGDP) and announced in February 2014. This markup uses different syntax to support the features of DPM, and adds a number of minor features. In general Robert Frank favored a terse syntax reminiscent of 1980s TROFF syntax. It would not be difficult to convert a DPM-marked text to one that is marked up with basic fpn using search and replace; for example chapter heads in DPM are marked with four newlines, and in fpn with a leading .h2.

fpgen

A bit earlier, in July 2013, the independent volunteers of PGDP Canada announced their own new markup style, fpgen. Documented in the DP-canada WIKI, fpgen is also the work of an "rfrank", in this case Roger Frank, who tended to favor an XML-like bracketed syntax. Again it would not be difficult to convert a DPM text to an fpgen one; for example a DPM chapter head marked with four newlines would become <heading level='1' id="ch01">Head Text</heading>

my-dpm (blush)

I am not immune to N.I.H. and the temptation to define markup syntax. In PPQT version 1 I supported a number of extensions to DPM, including right-aligned text and a syntax for tables. I designed these features based off of PPQT's model, Guiguts, which had its own simple extensions of DPM. For example, in DPM a block quote is

/Q
Quote text...
Q/

Guiguts extended this to allow specifying the first, left, and right indents so:

/Q[8,4,12]
Quote text with 8-char first indent, 4-char left indent, 
and 12-char right indent...
/Q

My version supported in PPQT V.1 allowed instead,

/Q F:8 L:4 R12
Quote text with 8-char first indent, 4-char left indent, 
and 12-char right indent...
Q/

Guiguts had a simple ASCII table markup; I extended it with additional syntax for column alignments and widths. I also added /R..R/ for right-aligned text and /C..C/ for centered text.

What to support with PPQT?

In a way, the choice of markup hardly matters, because the markup disappears before the book reaches its destination at Gutenberg.org. A marked-up document is a transient state between the original OCR text and the final etext/html/EPUB files. So the choice of markup is merely a convenience for the PPer. It is a way for her to encode decisions about how the book should be formatted: these lines are a poem, these lines are a table; this is emphasized text, etc.

However, the choice of markup is controlled by the software used for creating the final output. Robert Frank has a Python program to convert an fpn text to EPUB and HTML. Roger Frank of PGDP Canada has a, guess what, Python 3 program to convert fpgen to EPUB and HTML.

And both Guiguts and PPQT V.1 have code to convert DPM to HTML.

What should PPQT V.2 do? Should it contain code to convert fpn or fpgen to some other output format? Should it retain the V.1 HTML converter? Or should it be markup-agnostic?

Agnosticism

By markup-agnostic I mean, have only the features needed to make a clean job of finalizing an etext,including:

  • Support for image display alongside text,
  • proofer's notes saved in the metadata,
  • an extensive find/replace,
  • the character and word tables (so important for spell-check and finding other missed errors),
  • the Footnote panel with essential aids for cleaning and renumbering footnotes,
  • automatic calling of gutcheck or (better) bookloupe and a tabular display of the resulting diagnostics,

And just stop there, and say: ok, PPer, now you have a smooth DPM text, you can go on to use the editor and regex find/replace to convert this to any markup you like, and save the file, and process it using software from whomever.

Translators

Another option would be to offer automated translation from DPM to fpn and/or fpgen. And further, it would be possible to foist the job of coding those translators off on the people who want those markups. I've already floated this as an idea to DP Canada: that they could write the fpgen-erator to some API that I could provide.

Enter Pandoc

And then, there's Pandoc. Pandoc is a universal markup-translator. It reads texts in a variety of markup styles, and it writes output in an even wider variety including EPUB, LaTex, and PDF. It is widely used and widely praised, and in principle, could completely replace the programs written by both rfranks, generating any possible desired output format from code that is widely used and supported by an active community.

All that is needed is a way to get a post-processed etext into Pandoc. Unfortunately although Pandoc accepts a number of markups, that list does not include dpm, fpgen or fpn.

Pandoc does offer two general input formats. One is its own internal format represented in JSON format. The other is its own extended Markdown. This markup—let's call it pem for Pandoc's Extended Markdown—supports everything that dpm supports. It does not support quite everything that fpn supports, and I'm not sure about whether it is a superset of fpgen or not.

I can picture PPQT supporting a batch conversion of dpm to pem, for example as a command under the File menu: File > Save to Pandoc, and this as a replacement for the old HTML conversion step. This would convert a dpm text to a pem text and write it. However that wouldn't be much use to a PPer who didn't have access to Pandoc, because only Pandoc supports pem.

If I can work out how to distribute a Pandoc executable with PPQT in any platform, I can also imagine having automated "Save to Epub" and "Save to HTML" commands, that would generate a pem stream and feed it down a pipe to a Pandoc command, with the output to the designated file.

What About HTML?

PPQT V.1 has only one aid for HTML, the HTML Preview panel. When editing an HTML document, you can get a rendering of it in a QWebFrame. It's only a bit more convenient than saving the file and opening it in a separate browser (you avoid having to do a ctl-s in PPQT, click in the browser, click the reload button, then click back to PPQT to edit).

But should V.2 have any HTML support at all? The point in editing HTML inside PPQT was that the auto-converted HTML is awful to look at and needed a lot of hand-tweaking and customizing. But when HTML conversion is pushed off to an external program, whether that is Pandoc or an effort by one of the rfranks, is there any point in editing the resulting HTML output? Or is one supposed to do confine one's tweaking to the marked-up fp[ge]n file, and treat the HTML as a write-only output?

And supposing PPQT retained its HTML preview panel, should it also have a preview panel for EPUB? (Is that possible?) It would be kind of slick if, when you opened a file with an html suffix, you automatically got an HTML preview panel on the right, while if you opened a file with a .mobi suffix, you got an EPUB preview panel...

Also given HTML support of any kind, what about W3C Validation? Back when I assumed V.2 would have its own HTML conversion, I also speculated that it should support an automatic upload to the validation site, automatic download of the error list, and display of that in a panel such that you could click on a diagnostic and jump the editor to the referenced line. Is that still useful when HTML is being generated by an external program? Or is validation and W3C conformity now the responsibility of the external program?

I welcome the comments of any of my readers on these issues.