Advanced settings
The Advanced tab controls what happens to the text: getting the full article, cleaning the HTML, translating and rewriting it, filling custom fields, and how the feed is downloaded.
Click to enlargeExtract full-text articles
Many feeds contain only a short summary. With this option, RSS Retriever opens the link of every item and extracts the complete article from the source page.
- Don’t extract (default)
- Use the content of the feed as it is.
- Use Full-Text RSS script
- Available when the Full-Text RSS script is installed (see General Settings). The script detects the article on most news and blog pages automatically, so it also works for aggregator feeds that link to many different sites. It may miss parts such as embedded videos. If the page cannot be retrieved directly, RSS Retriever tries the Wayback Machine when Use Wayback Machine is on.
- Use custom settings
- You tell RSS Retriever which HTML element holds the article. Precise, but it works only for one site at a time, because every site has its own layout. Not suitable for aggregator feeds such as Google News.
Custom extraction settings
- Container tag
- The HTML tag that wraps the article on the source pages, such as
article,divorsection. Default:div. - Attributes (JSON format)
- Attribute-value pairs that identify the right element when there are several tags of the same name, in JSON:
{"class": "article-body"}or{"class": "article", "id": "main"}. - Inclusive
- Tick to keep the container tag itself with its attributes in the result. Leave it off to get only the inner content.
To find the right element, open an article of the source site in your browser, right-click the text and choose Inspect. Look for the element that contains the whole article text and nothing else. Check the result with the Full-text article preview.
When extraction is on and the article cannot be retrieved, the item is skipped and the log says “Unable to retrieve full-text content”. RSS Retriever never publishes the short version instead.
HTML tags to strip
A comma-separated list of tags removed from the posts, for example h1, img, script, style. Only the tags are removed; the text inside them stays.
Remove outer HTML elements
Removes whole HTML blocks, including everything inside them. Write a semicolon-separated list; each rule is a tag name followed by its attribute-value pairs in JSON format:
div {"class": "share-box"}; p {"class": "description", "id": "block"}
The first rule removes every <div class="share-box"> with its share buttons; the second removes every <p> that has both class="description" and id="block". Write tag names and JSON exactly, or the rule will not match.
Convert Markdown to HTML
Some AI models return Markdown (### Heading, **bold**, * item) even when asked for HTML. This option converts it into HTML. It applies only to the text produced by AI shortcodes. The converter library is downloaded automatically the first time it is needed.
UTF-8 encoding
Converts an ISO-8859-1 string to UTF-8. Use it for feeds that contain invalid UTF-8 bytes (such as <0x92>) and fail to parse otherwise.
Convert character encoding
For feeds delivered in a national character set such as windows-1251 or ISO-8859-1: the feed is converted to UTF-8.
Sanitize content
Sanitizes the post content: validates UTF-8, converts lone < characters to entities, strips unsafe tags such as iframe, embed, style and script, and removes tabs, line breaks and extra white space.
Balance HTML tags
Closes unclosed tags and removes stray closing tags in the content and the excerpt with the WordPress tag balancer. Helpful for feeds with broken markup or shortened HTML. Enable it only when posts have a reasonable size: very long content can use a lot of memory.
Shorten post excerpts
The maximum number of words left in the excerpt. 0 removes excerpts completely; an empty field keeps them as they are.
Link handling
- Keep links intact (default).
- Remove all links: the link text stays, the links go.
- Remove all links except links in images.
- Remove links from images only.
To change link attributes on the fly instead (nofollow, new window), use the runtime options.
Translation
Translates the title, content and excerpt of every post. The HTML structure is kept.
- Do not translate (default)
- AI Translate
- Choose the Target language and the AI engine (default
openai-gpt-4o-mini, a fast and cheap choice). Works with any AI engine and supports many more languages than classic services. - Use DeepL Translate
- Choose the Target language. Tick Use DeepL API Free if your key belongs to the free DeepL API plan (up to 500,000 characters a month); untick it for DeepL API Pro keys.
- Use Yandex Translate
- Choose the Direction, for example English to Russian.
- Use Google Translate
- Choose the Source and Target languages.
Each service needs its API key on the Accounts page. If the translation fails, the post is not added. For multilingual sites, see Multilingual sites.
Content filtering checks the item before it is translated, so write your keywords in the language of the source. Items that do not pass are never sent to the paid translation service.
Built-in synonymizer
Disable, Use built-in synonymizer before content spinner or Use built-in synonymizer after content spinner (default). The synonym table is edited on the Synonymizer/Rewriter page; with an empty table nothing changes.
Content spinner
Rewrites the post with an external service, so the text differs from the source:
- Disable (default)
- AI Spinner
- Rewrites the text with the AI engine you choose (default
openai-gpt-4o-mini), keeps the HTML structure and works with large articles and in many languages, such as German, Spanish, French, Turkish or Chinese. The rewriting instructions are maintained on the CyberSEO server. AI providers refuse some topics; such articles cannot be spun. - SpinnerChief
- Uses your SpinnerChief API key and developer key (English).
- SpinRewriter
- Options: Protected terms (one term per line that must not be changed), Auto protected terms (protect capitalized words), Confidence level (low, medium, high), Auto sentences, Auto new paragraphs, Auto sentence trees, Use only synonyms, Text with Spintax and Nested Spintax.
- WordAi
- Options: Uniqueness (more conservative, regular, more adventurous), Return rewrites, Protect words and Use custom synonyms (from your WordAi account settings), and Avoid AI detection, which returns a version of the text that avoids AI detectors and ignores the other options.
Don’t synonymize titles
By default the synonymizer changes titles and content. Tick this option to keep titles unchanged.
Custom fields
Fills custom fields (post meta) of every post. One rule per line, in the format source->field_name.
Fixed values
Text in double quotes. Spintax is allowed.
"T-shirt"->product_type
"{red|blue|green}"->color
XML tag values
The name of a tag of the feed item. For this item:
<price currency="USD">29.99</price>
<isbn>978-3-16-148410-0</isbn>
these rules write the price into WooCommerce fields and the ISBN into its own field:
price->_price
price->_regular_price
isbn->isbn
XML attributes
Tag name, colon, attribute name:
price:currency->price_currency
Content of the post: HTML elements
A tag name with attributes in JSON format takes the value of that element from the generated post. For <span class="price">129.95</span><span class="old-price">249.95</span>:
span {"class": "price"}->_price
span {"class": "old-price"}->_regular_price
Content of the post: regular expressions
Prefix the expression with regex:; the first captured group is stored. To store the URL of a featured image found in the content:
regex:<img class="featured" src="(.*?)"->thumb
Writing to post fields
Instead of a custom field name, the target can be one of %post_title%, %post_content%, %post_excerpt%, %post_tags% or %categories%. For example, take the title from a tag of a product feed, or the tags from a comma-separated tag:
product_name->%post_title%
keywords->%post_tags%
Values set here can be used in templates with %custom_fields[name]%. The special field thumb can be used as the source of the featured image, and video as a video to embed (see Media handling). Invalid rules are reported in the log as “Error in custom field rule”.
Connection settings
These options decide how the feed and its pages are downloaded. They are also available in Quick override default settings before you add a feed.
Proxy mode
No proxy (default) or Use proxy list: requests go through the proxies of the Proxy list. Use proxies for sources that block or limit many requests from one IP address.
Use Wayback Machine
Enabled by default. When the source blocks direct access (bot protection, Cloudflare, geographic restrictions, IP limits), RSS Retriever requests the latest archived copy from the Internet Archive’s Wayback Machine instead. This works for the feed, full-text articles and images.
The archive must already have a copy of the URL; very new articles and images may not be there yet. Archive requests take a little longer.
User agent
The browser identification sent with requests. If a feed fails with an error such as “Unable to acquire”, the site may block scripts; try the user agent of a browser or a crawler:
Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:128.0) Gecko/20100101 Firefox/128.0
Googlebot/2.1 (+http://www.google.com/bot.html)
FeedValidator/1.3
HTTP referrer
The Referer header sent with requests. Default: self, which sends the URL of the source itself. Some sites deliver content only to visitors coming from a certain page. Referer spoofing can violate the terms of a site; use it only when necessary.
HTTP headers
Extra request headers, one per line, as Name: value:
Accept-Language: en-US,en;q=0.9
Cookie: consent=yes