Files
ttrss-plugin-af-fulltext/lib/ExtractResult.php
Michal 9bc854ec44 af_fulltext: rule-driven extraction with a pluggable renderer
Replaces subscribing feeds through a self-hosted Full-Text RSS proxy. Feed URLs
go back to being real feed URLs and extraction happens inside tt-rss, using the
same ftr-site-config rules the proxy used.

The engine in lib/ has no tt-rss dependencies, so rules can be developed and
audited from the command line; init.php is a thin adapter over it.

Two findings from measuring the real subscription first, both of which shaped
the design:

- Firecrawl's own onlyMainContent is far too coarse to extract with (73KB of
  chrome on a Cloudflare post), but it is an excellent renderer. So it is used
  for rawHtml only and the rule engine does the extraction.
- A body rule that stops matching after a redesign falls through to Readability
  and still produces a plausible article, so the breakage is invisible. Every
  extraction now records which rule matched and whether it fell back; auditing
  the 34 live feeds surfaced five community rules that match nothing.

Custom rules included for the sites that needed them, including three comics
where the article is an image and text-scoring extractors return the wrong thing
or nothing at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011sSJdftQx5bW5HZHgUKF3i
2026-08-25 00:16:16 +01:00

58 lines
1.7 KiB
PHP

<?php
namespace AfFulltext;
/**
* The outcome of one extraction, diagnostics included.
*
* The diagnostics are the point. A site redesign turns a `body:` rule into a
* silent no-op that falls through to Readability, and the article still looks
* plausible -- that is how a broken webcomic rule went unnoticed here for years.
* Every consumer of this class can see exactly which rule ran and whether it
* actually matched anything.
*/
final class ExtractResult {
public string $html = '';
public ?string $title = null;
public ?string $author = null;
public ?string $date = null;
public string $backend = '';
public string $effective_url = '';
/** Rule files that contributed, nearest first. @var string[] */
public array $rule_sources = [];
/** True when a `body:` rule existed and selected at least one node. */
public bool $rule_matched = false;
/** True when the content came from Readability rather than a rule. */
public bool $fell_back = false;
/** True when a rule existed but matched nothing -- the stale-rule signal. */
public bool $rule_stale = false;
public int $pages = 1;
public int $timing_ms = 0;
/** @var string[] */
public array $errors = [];
public function ok(): bool {
return $this->html !== '';
}
/** One-line summary for logs and the prefs UI. */
public function summary(): string {
$rule = $this->rule_sources ? basename($this->rule_sources[0]) : 'none';
$how = match (true) {
$this->rule_matched => "rule=$rule",
$this->rule_stale => "rule=$rule STALE->readability",
default => 'readability',
};
return sprintf('%s %s %db %dms%s', $this->backend, $how, strlen($this->html),
$this->timing_ms, $this->pages > 1 ? " pages={$this->pages}" : '');
}
}