Link-roundup themes, press-clipping posts, and “read the original” CTAs all need the first URL inside post HTML. The same scan can promote the first <img> to a featured image when an editor forgets to set one. In both cases you parse post_content (or filtered the_content) and pull attributes from the DOM.
Regex feels faster to write. It fails on real WordPress HTML. Prefer PHP’s DOMDocument, suppress libxml noise, escape every URL you print, and cache results when you run this on archives.
For custom theme and plugin work: WordPress developer.
Method 1: first link with DOMDocument
function wppoland_get_first_link_url( $content ) {
if ( empty( $content ) ) {
return false;
}
$doc = new DOMDocument();
libxml_use_internal_errors( true );
$wrapped = '<?xml encoding="utf-8" ?>' . $content;
$doc->loadHTML( $wrapped, LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD );
libxml_clear_errors();
$links = $doc->getElementsByTagName( 'a' );
if ( $links->length > 0 ) {
$href = $links->item( 0 )->getAttribute( 'href' );
return $href !== '' ? $href : false;
}
return false;
}LIBXML_HTML_NOIMPLIED is optional and behaves differently across libxml versions. If your host’s libxml is old and fragments break, drop those flags and keep the UTF-8 wrapper plus libxml_use_internal_errors( true ).
Usage in the Loop
$link = wppoland_get_first_link_url( get_the_content( null, false ) );
if ( $link ) {
printf(
'<a href="%s" class="read-more-external" rel="noopener noreferrer">Read original</a>',
esc_url( $link )
);
}Use get_post()->post_content when you need raw blocks without the_content filters (shortcodes, embeds). Use get_the_content() when you want the same HTML the front end would render.
Method 2: first image URL
function wppoland_get_first_image_url( $content ) {
if ( empty( $content ) ) {
return false;
}
$doc = new DOMDocument();
libxml_use_internal_errors( true );
$doc->loadHTML( '<?xml encoding="utf-8" ?>' . $content );
libxml_clear_errors();
$images = $doc->getElementsByTagName( 'img' );
if ( $images->length > 0 ) {
$src = $images->item( 0 )->getAttribute( 'src' );
return $src !== '' ? $src : false;
}
return false;
}Auto-set featured image on save
add_action( 'save_post_post', function( $post_id ) {
if ( wp_is_post_revision( $post_id ) || wp_is_post_autosave( $post_id ) ) {
return;
}
if ( has_post_thumbnail( $post_id ) ) {
return;
}
$post = get_post( $post_id );
if ( ! $post ) {
return;
}
$image_url = wppoland_get_first_image_url( $post->post_content );
if ( ! $image_url ) {
return;
}
$attachment_id = attachment_url_to_postid( $image_url );
if ( $attachment_id ) {
set_post_thumbnail( $post_id, $attachment_id );
}
}, 20 );attachment_url_to_postid() only resolves media already in the library. External hotlinked images will not become thumbnails without a download step - skip that unless you own the rights and have a clear import path.
Method 3: regex (reference only)
function wppoland_get_first_link_url_regex( $content ) {
if ( empty( $content ) ) {
return false;
}
if ( preg_match( '/<a\s[^>]*href=["\']([^"\']+)["\'][^>]*>/i', $content, $matches ) ) {
return $matches[1];
}
return false;
}Why not ship this in production:
- Attribute order and single vs double quotes still surprise people.
- Nested markup and Gutenberg figure/link wrappers confuse naive patterns.
- Harder to extend to
title,rel, or “first external only”.
Keep regex for one-off CLI cleanups, not theme templates.
Advanced: full data from the first anchor
function wppoland_get_first_link_data( $content ) {
if ( empty( $content ) ) {
return false;
}
$doc = new DOMDocument();
libxml_use_internal_errors( true );
$doc->loadHTML( '<?xml encoding="utf-8" ?>' . $content );
libxml_clear_errors();
$links = $doc->getElementsByTagName( 'a' );
if ( 0 === $links->length ) {
return false;
}
$node = $links->item( 0 );
return array(
'url' => $node->getAttribute( 'href' ),
'text' => $node->textContent,
'title' => $node->getAttribute( 'title' ),
'target' => $node->getAttribute( 'target' ),
'rel' => $node->getAttribute( 'rel' ),
'class' => $node->getAttribute( 'class' ),
);
}$data = wppoland_get_first_link_data( get_the_content( null, false ) );
if ( $data && $data['url'] ) {
printf(
'<a href="%s" title="%s" target="%s" rel="%s">%s</a>',
esc_url( $data['url'] ),
esc_attr( $data['title'] ),
esc_attr( $data['target'] ? $data['target'] : '_blank' ),
esc_attr( $data['rel'] ? $data['rel'] : 'noopener noreferrer' ),
esc_html( $data['text'] )
);
}First external link only
Roundup posts often start with an internal “jump to comments” or related-post link. Skip those:
function wppoland_get_first_external_link( $content ) {
if ( empty( $content ) ) {
return false;
}
$doc = new DOMDocument();
libxml_use_internal_errors( true );
$doc->loadHTML( '<?xml encoding="utf-8" ?>' . $content );
libxml_clear_errors();
$site_host = wp_parse_url( home_url(), PHP_URL_HOST );
$links = $doc->getElementsByTagName( 'a' );
foreach ( $links as $link ) {
$href = trim( $link->getAttribute( 'href' ) );
if ( '' === $href || '#' === $href[0] ) {
continue;
}
$host = wp_parse_url( $href, PHP_URL_HOST );
if ( $host && $host !== $site_host ) {
return $href;
}
}
return false;
}Caching so archives stay cheap
Parsing HTML for every card on a category archive is wasteful. Store on save:
function wppoland_cache_first_link( $post_id ) {
if ( wp_is_post_revision( $post_id ) || wp_is_post_autosave( $post_id ) ) {
return;
}
$post = get_post( $post_id );
if ( ! $post || 'post' !== $post->post_type ) {
return;
}
$url = wppoland_get_first_link_url( $post->post_content );
update_post_meta( $post_id, '_wppoland_first_link', $url ? $url : '' );
}
add_action( 'save_post_post', 'wppoland_cache_first_link', 30 );
function wppoland_get_cached_first_link( $post_id = null ) {
$post_id = $post_id ?: get_the_ID();
$cached = get_post_meta( $post_id, '_wppoland_first_link', true );
if ( '' === $cached ) {
return false;
}
if ( $cached ) {
return $cached;
}
// Backfill for old posts.
$post = get_post( $post_id );
$url = $post ? wppoland_get_first_link_url( $post->post_content ) : false;
update_post_meta( $post_id, '_wppoland_first_link', $url ? $url : '' );
return $url;
}Transients work too; post meta survives cache flushes and is easier to invalidate on save_post.
Security and edge cases
- Always
esc_url()on output. Never trust content authors blindly. - Empty
href,javascript:URLs, anddata:URIs should be rejected if you only want http(s). - Block markup may wrap links in
<figure>or use relative paths -DOMDocumentstill finds the<a>; resolve relative URLs withwp_make_link_relative/ site URL as needed. - Run extraction in admin or on save when possible so front-end PHP stays thin.
Gutenberg, classic content, and filtered HTML
post_content in the database stores blocks as HTML comments plus inner HTML. DOMDocument still finds <a> and <img> nodes inside those blocks. If you run apply_filters( 'the_content', $post->post_content ) first, shortcodes and embeds expand - which can insert extra links (share buttons, related posts injected by a plugin). For “first editorial link” semantics, parse raw post_content. For “first link the visitor would see”, parse after the_content.
Classic editor posts and imported XML often contain unclosed tags. That is exactly where regex fails and DOMDocument + libxml_use_internal_errors( true ) earns its keep.
Reject dangerous schemes
Before you echo a URL as href, allow only http(s) (and optionally relative paths you resolve yourself):
function wppoland_safe_http_url( $url ) {
$url = trim( $url );
if ( '' === $url ) {
return false;
}
$parsed = wp_parse_url( $url );
if ( empty( $parsed['scheme'] ) ) {
// Site-relative path.
return esc_url_raw( home_url( $url ) );
}
$scheme = strtolower( $parsed['scheme'] );
if ( ! in_array( $scheme, array( 'http', 'https' ), true ) ) {
return false;
}
return esc_url_raw( $url );
}Drop javascript:, data:, and vbscript: early. Editors paste weird things; treat content as untrusted input.
Unit-test the helper
A minimal PHPUnit / WP integration assertion saves regressions when someone “improves” the regex later:
public function test_first_link_skips_hash_and_finds_external() {
$html = '<p><a href="#comments">Comments</a> <a href="https://example.com/x">X</a></p>';
$this->assertSame(
'https://example.com/x',
wppoland_get_first_external_link( $html )
);
}Fixture a post with UTF-8 link text (Norwegian ø, Polish ł) and confirm textContent is not mojibake after loadHTML.
CLI backfill for old posts
After deploying post-meta caching, backfill once:
wp post list --post_type=post --format=ids | xargs -n1 -I% wp eval 'wppoland_cache_first_link( % );'Or a small WP_CLI command that batches by ID range so you do not load the entire table into memory on large sites.
Summary
- Prefer
DOMDocumentover regex for first-link and first-image extraction. - Suppress libxml errors; preserve UTF-8 when loading HTML.
- Escape every attribute you print; add
rel="noopener noreferrer"for_blank. - Cache on
save_postfor archive performance. - Filter for external hosts when building aggregator-style CTAs.
Need a production-safe helper packaged as a must-use plugin? Contact with the theme type (classic vs block) and whether you need first link, first image, or both.






