Advanced settings

The Advanced tab controls what happens to the text: getting the full article, cleaning the HTML, translating and rewriting it, filling custom fields, and how the feed is downloaded.

The Advanced tab with extraction, translation and connection settingsClick to enlarge

Extract full-text articles

Many feeds contain only a short summary. With this option, RSS Retriever opens the link of every item and extracts the complete article from the source page.

Don’t extract (default)
Use the content of the feed as it is.
Use Full-Text RSS script
Available when the Full-Text RSS script is installed (see General Settings). The script detects the article on most news and blog pages automatically, so it also works for aggregator feeds that link to many different sites. It may miss parts such as embedded videos. If the page cannot be retrieved directly, RSS Retriever tries the Wayback Machine when Use Wayback Machine is on.
Use custom settings
You tell RSS Retriever which HTML element holds the article. Precise, but it works only for one site at a time, because every site has its own layout. Not suitable for aggregator feeds such as Google News.

Custom extraction settings

Container tag
The HTML tag that wraps the article on the source pages, such as article, div or section. Default: div.
Attributes (JSON format)
Attribute-value pairs that identify the right element when there are several tags of the same name, in JSON: {"class": "article-body"} or {"class": "article", "id": "main"}.
Inclusive
Tick to keep the container tag itself with its attributes in the result. Leave it off to get only the inner content.

To find the right element, open an article of the source site in your browser, right-click the text and choose Inspect. Look for the element that contains the whole article text and nothing else. Check the result with the Full-text article preview.

No full text, no post

When extraction is on and the article cannot be retrieved, the item is skipped and the log says “Unable to retrieve full-text content”. RSS Retriever never publishes the short version instead.

HTML tags to strip

A comma-separated list of tags removed from the posts, for example h1, img, script, style. Only the tags are removed; the text inside them stays.

Remove outer HTML elements

Removes whole HTML blocks, including everything inside them. Write a semicolon-separated list; each rule is a tag name followed by its attribute-value pairs in JSON format:

div {"class": "share-box"}; p {"class": "description", "id": "block"}

The first rule removes every <div class="share-box"> with its share buttons; the second removes every <p> that has both class="description" and id="block". Write tag names and JSON exactly, or the rule will not match.

Convert Markdown to HTML

Some AI models return Markdown (### Heading, **bold**, * item) even when asked for HTML. This option converts it into HTML. It applies only to the text produced by AI shortcodes. The converter library is downloaded automatically the first time it is needed.

UTF-8 encoding

Converts an ISO-8859-1 string to UTF-8. Use it for feeds that contain invalid UTF-8 bytes (such as <0x92>) and fail to parse otherwise.

Convert character encoding

For feeds delivered in a national character set such as windows-1251 or ISO-8859-1: the feed is converted to UTF-8.

Sanitize content

Sanitizes the post content: validates UTF-8, converts lone < characters to entities, strips unsafe tags such as iframe, embed, style and script, and removes tabs, line breaks and extra white space.

Balance HTML tags

Closes unclosed tags and removes stray closing tags in the content and the excerpt with the WordPress tag balancer. Helpful for feeds with broken markup or shortened HTML. Enable it only when posts have a reasonable size: very long content can use a lot of memory.

Shorten post excerpts

The maximum number of words left in the excerpt. 0 removes excerpts completely; an empty field keeps them as they are.

To change link attributes on the fly instead (nofollow, new window), use the runtime options.

Translation

Translates the title, content and excerpt of every post. The HTML structure is kept.

Do not translate (default)
 
AI Translate
Choose the Target language and the AI engine (default openai-gpt-4o-mini, a fast and cheap choice). Works with any AI engine and supports many more languages than classic services.
Use DeepL Translate
Choose the Target language. Tick Use DeepL API Free if your key belongs to the free DeepL API plan (up to 500,000 characters a month); untick it for DeepL API Pro keys.
Use Yandex Translate
Choose the Direction, for example English to Russian.
Use Google Translate
Choose the Source and Target languages.

Each service needs its API key on the Accounts page. If the translation fails, the post is not added. For multilingual sites, see Multilingual sites.

Filters run before translation

Content filtering checks the item before it is translated, so write your keywords in the language of the source. Items that do not pass are never sent to the paid translation service.

Built-in synonymizer

Disable, Use built-in synonymizer before content spinner or Use built-in synonymizer after content spinner (default). The synonym table is edited on the Synonymizer/Rewriter page; with an empty table nothing changes.

Content spinner

Rewrites the post with an external service, so the text differs from the source:

Disable (default)
 
AI Spinner
Rewrites the text with the AI engine you choose (default openai-gpt-4o-mini), keeps the HTML structure and works with large articles and in many languages, such as German, Spanish, French, Turkish or Chinese. The rewriting instructions are maintained on the CyberSEO server. AI providers refuse some topics; such articles cannot be spun.
SpinnerChief
Uses your SpinnerChief API key and developer key (English).
SpinRewriter
Options: Protected terms (one term per line that must not be changed), Auto protected terms (protect capitalized words), Confidence level (low, medium, high), Auto sentences, Auto new paragraphs, Auto sentence trees, Use only synonyms, Text with Spintax and Nested Spintax.
WordAi
Options: Uniqueness (more conservative, regular, more adventurous), Return rewrites, Protect words and Use custom synonyms (from your WordAi account settings), and Avoid AI detection, which returns a version of the text that avoids AI detectors and ignores the other options.

Don’t synonymize titles

By default the synonymizer changes titles and content. Tick this option to keep titles unchanged.

Custom fields

Fills custom fields (post meta) of every post. One rule per line, in the format source->field_name.

Fixed values

Text in double quotes. Spintax is allowed.

"T-shirt"->product_type
"{red|blue|green}"->color

XML tag values

The name of a tag of the feed item. For this item:

<price currency="USD">29.99</price>
<isbn>978-3-16-148410-0</isbn>

these rules write the price into WooCommerce fields and the ISBN into its own field:

price->_price
price->_regular_price
isbn->isbn

XML attributes

Tag name, colon, attribute name:

price:currency->price_currency

Content of the post: HTML elements

A tag name with attributes in JSON format takes the value of that element from the generated post. For <span class="price">129.95</span><span class="old-price">249.95</span>:

span {"class": "price"}->_price
span {"class": "old-price"}->_regular_price

Content of the post: regular expressions

Prefix the expression with regex:; the first captured group is stored. To store the URL of a featured image found in the content:

regex:<img class="featured" src="(.*?)"->thumb

Writing to post fields

Instead of a custom field name, the target can be one of %post_title%, %post_content%, %post_excerpt%, %post_tags% or %categories%. For example, take the title from a tag of a product feed, or the tags from a comma-separated tag:

product_name->%post_title%
keywords->%post_tags%

Values set here can be used in templates with %custom_fields[name]%. The special field thumb can be used as the source of the featured image, and video as a video to embed (see Media handling). Invalid rules are reported in the log as “Error in custom field rule”.

Connection settings

These options decide how the feed and its pages are downloaded. They are also available in Quick override default settings before you add a feed.

Proxy mode

No proxy (default) or Use proxy list: requests go through the proxies of the Proxy list. Use proxies for sources that block or limit many requests from one IP address.

Use Wayback Machine

Enabled by default. When the source blocks direct access (bot protection, Cloudflare, geographic restrictions, IP limits), RSS Retriever requests the latest archived copy from the Internet Archive’s Wayback Machine instead. This works for the feed, full-text articles and images.

The archive must already have a copy of the URL; very new articles and images may not be there yet. Archive requests take a little longer.

User agent

The browser identification sent with requests. If a feed fails with an error such as “Unable to acquire”, the site may block scripts; try the user agent of a browser or a crawler:

Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:128.0) Gecko/20100101 Firefox/128.0
Googlebot/2.1 (+http://www.google.com/bot.html)
FeedValidator/1.3

HTTP referrer

The Referer header sent with requests. Default: self, which sends the URL of the source itself. Some sites deliver content only to visitors coming from a certain page. Referer spoofing can violate the terms of a site; use it only when necessary.

HTTP headers

Extra request headers, one per line, as Name: value:

Accept-Language: en-US,en;q=0.9
Cookie: consent=yes