I can retrieve the text of a web page, let's say https://*.com/questions with some real and made up links:
我可以检索一个网页的文本,让我们说https://*.com/questions与一些真实和组成的链接:
/questions /tags /questions?sort=votes /questions?sort=active randompage.aspx ../coolhomepage.aspx
Knowing my originating page was https://*.com/questions is there a way in .Net to resolve the links to this?
知道我的原始页面是https://*.com/questions有没有办法在.Net解决这个链接?
https://*.com/questions https://*.com/tags https://*.com/questions?sort=votes https://*.com/questions?sort=active https://*.com/questions/randompage.aspx https://*.com/coolhomepage.aspx
Kind of like the way a Browser is smart enough to resolve the links.
有点像浏览器足够聪明地解析链接的方式。
=========================== Update - Using David's solution:
===========================更新 - 使用David的解决方案:
'Regex to match all <a ... /a> links Dim myRegEx As New Regex("\<\s*a (?# Find opening <a tag) " & _ ".+?href\s*=\s*['""] (?# Then all to href=' or "" ) " & _ "(?<href>.*?)['""] (?# Then all to the next ' or "" ) " & _ ".*?\> (?# Then all to > ) " & _ "(?<name>.*?)\<\s*/a\s*\> (?# Then all to </a> ) ", _ RegexOptions.IgnoreCase Or _ RegexOptions.IgnorePatternWhitespace Or _ RegexOptions.Multiline) 'MatchCollection to hold all the links that are matched Dim myMatchCollection As MatchCollection myMatchCollection = myRegEx.Matches(Me._RawPageText) 'Loop through all matches and evaluate the value of the href attribute. For i As Integer = 0 To myMatchCollection.Count - 1 Dim thisLink As String = "" thisLink = myMatchCollection(i).Groups("href").Value() 'This checks for Javascript and Mailto links. 'This is not complete. There are others to check I just haven't encountered them yet. If thisLink.ToLower.StartsWith("javascript") Then thisLink = "JAVASCRIPT: " & thisLink ElseIf thisLink.ToLower.StartsWith("mailto") Then thisLink = "MAILTO: " & thisLink Else Dim baseUri As New Uri(Me.URL) If Not thisLink.ToLower.StartsWith("http") Then 'This is a partial URL so we will assume that it's relative to our originating URL Dim myUri As New Uri(baseUri, thisLink) thisLink = "RELATIVE LOCAL LINK: RESOLVED: " & myUri.ToString() & " ORIGINAL: " & thisLink Else 'The link starts with HTTP, determine if part of base host or is outside host. Dim ThisUri As New Uri(thisLink) If ThisUri.Host.ToLower = baseUri.Host.ToLower Then thisLink = "INSIDE COMPLETE LINK: " & thisLink Else thisLink = "OUTSIDE LINK: " & thisLink End If End If End If 'I'm storing the found links into a Generic.List(Of String) 'This link has descriptive text added to it. 'TODO: Make collection to hold only unique internal links. Me._Links.Add(thisLink) Next
3 个解决方案
#1
You mean like this?
你的意思是这样的?
Uri baseUri = new Uri("http://www.contoso.com");
Uri myUri = new Uri(baseUri, "catalog/shownew.htm");
Console.WriteLine(myUri.ToString());
Sample comes from http://msdn.microsoft.com/en-us/library/9hst1w91.aspx
示例来自http://msdn.microsoft.com/en-us/library/9hst1w91.aspx
#2
If you mean server-side, you can use ResolveUrl()
:
如果您的意思是服务器端,则可以使用ResolveUrl():
string url = ResolveUrl("~/questions");
#3
I dont understand what you mean by "resolve" in this context, but you can try inserting a base html element. Since you asked how the browser would handle it.
我不明白你在这个上下文中的“解决”是什么意思,但你可以尝试插入一个基本的html元素。既然你问过浏览器将如何处理它。
"The <base>
tag specifies a default address or a default target for all links on a page."
“
#1
You mean like this?
你的意思是这样的?
Uri baseUri = new Uri("http://www.contoso.com");
Uri myUri = new Uri(baseUri, "catalog/shownew.htm");
Console.WriteLine(myUri.ToString());
Sample comes from http://msdn.microsoft.com/en-us/library/9hst1w91.aspx
示例来自http://msdn.microsoft.com/en-us/library/9hst1w91.aspx
#2
If you mean server-side, you can use ResolveUrl()
:
如果您的意思是服务器端,则可以使用ResolveUrl():
string url = ResolveUrl("~/questions");
#3
I dont understand what you mean by "resolve" in this context, but you can try inserting a base html element. Since you asked how the browser would handle it.
我不明白你在这个上下文中的“解决”是什么意思,但你可以尝试插入一个基本的html元素。既然你问过浏览器将如何处理它。
"The <base>
tag specifies a default address or a default target for all links on a page."
“